Quick answer
Start with the PIRB leaderboard: 41 Polish retrieval tasks put together by a team at OPI PIB, Poland's National Information Processing Institute. As of 30 September 2026, PolDense-1B leads with 64.11 NDCG@10, ahead of models with 8-9 billion parameters. Popular multilingual models do worse in Polish: Qwen3-Embedding-8B scores 60.93, multilingual-e5-large 57.29 and BGE-M3 55.98. The biggest gain comes from reranking: the Polish polish-reranker-roberta-v3 (443M) reaches 65.17, close to Qwen3-Reranker-8B. Use the leaderboard to get down to two or three candidates, then let a test on about a hundred questions from your own documents decide.
Why model choice matters more in Polish
An embedding model turns a passage of text into a vector, and retrieval finds the passages whose vectors sit closest to the question's vector. If the model doesn't understand the language of your documents well, even the best language model at the end of the chain gets the wrong passages. We explain the whole chain in how RAG makes AI smarter.
Polish makes this harder in a few ways:
- Inflection. “Umowa”, “umowy”, “umowie”, “umową” are all forms of the word for “contract”. Full-text search without a Polish analyser treats them as different words.
- Less training data. Multilingual models learn mostly from English text. Polish gets a sliver of the data, and specialist vocabulary (law, medicine, public administration) gets even less.
- Mixed documents. Polish companies mix Polish and English: bilingual contracts, technical documentation, correspondence. The model has to find a Polish passage for an English question and the other way round.
English benchmarks (such as English MTEB) say little about any of this, which is why Polish leaderboards are worth a look.
What PIRB is and how to read it
PIRB (Polish Information Retrieval Benchmark) was published in 2024 by Sławomir Dadas, Michał Perełkiewicz and Rafał Poświata of OPI PIB (Dadas et al., 2024). It covers 41 retrieval tasks in five groups:
- MaupQA: an automatically created Polish question answering dataset (from IPI PAN);
- PolEval-2022: tasks from the PolEval challenge, including legal questions;
- BEIR-PL: BEIR datasets machine-translated into Polish (by Wrocław University of Science and Technology);
- Web Datasets: real questions and answers from Polish websites, including legal and medical ones;
- Other: including MFAQ and GPT-exams.
The main metric is NDCG@10, which scores whether the right passages appear in the top ten results and how high. The leaderboard has three sections: standalone models, hybrid retrieval, and two-stage retrieval with a reranker.
For a broader view there's PL-MTEB from the same authors: 30 tasks in five categories (classification, clustering, pair classification, retrieval and semantic similarity), merged into MTEB (Poświata et al., 2024). For RAG, retrieval is what matters most, and PIRB covers it in the most depth.
One note for completeness: PIRB, PolDense, the mmlw family and the Polish rerankers all come from the same team. The benchmark is open and the evaluation code is public (github.com/sdadas/pirb), but it's one more reason to confirm leaderboard results on your own data.
Embedding model leaderboard (as of 30 September 2026)
Average NDCG@10 across 41 tasks, computed from the PIRB leaderboard data. The “legal” column averages four legal tasks: two from PolEval-2022 (legal questions) plus eprawnik and specprawnik.
| Model | Parameters | PIRB average | PIRB legal | Licence |
|---|---|---|---|---|
| OPI-PIB/PolDense-1B | 1.0B | 64.11 | 73.05 | Gemma |
| nvidia/llama-embed-nemotron-8b | 7.5B | 63.73 | 69.39 | non-commercial only |
| BAAI/bge-multilingual-gemma2 | 9.2B | 63.26 | 71.82 | Gemma |
| OPI-PIB/PolDense-400M | 395M | 63.21 | 71.91 | Gemma |
| sdadas/stella-pl-retrieval-8k | 1.5B | 62.69 | 68.38 | Gemma |
| OPI-PIB/PolDense-150M | 149M | 61.07 | 68.54 | Gemma |
| Qwen/Qwen3-Embedding-8B | 7.6B | 60.93 | 64.49 | Apache 2.0 |
| Qwen/Qwen3-Embedding-4B | 4.0B | 59.19 | 62.38 | Apache 2.0 |
| sdadas/mmlw-retrieval-roberta-large | 435M | 58.46 | 60.13 | Apache 2.0 |
| intfloat/multilingual-e5-large | 560M | 57.29 | 57.89 | MIT |
| BAAI/bge-m3 (dense) | 568M | 55.98 | 59.80 | MIT |
| Qwen/Qwen3-Embedding-0.6B | 596M | 50.71 | 50.80 | Apache 2.0 |
| BM25 (no model) | n/a | 41.85 | 45.24 | n/a |
What this tells you:
- Size isn't everything. PolDense-400M scores close to models about 20 times its size. According to its model card, PolDense-150M comes close to Qwen3-Embedding-8B (model card), and OPI PIB says the smallest variants (17M, 32M, 68M) are designed to run on CPUs and edge devices (OPI PIB).
- Multilingual models lose ground in Polish. Qwen3-Embedding-8B ranked first on multilingual MTEB in June 2025 (70.58) (model card). In Polish it's solid, but behind models trained for Polish. The small 0.6B version does worse than multilingual-e5-large.
- Domain widens the gaps. On specprawnik (questions from a Polish legal website), PolDense-1B scores 46.6, Qwen3-Embedding-8B 31.3 and BGE-M3 25.5. If your documents are specialised, the overall average may understate the difference.
- Licences rule out candidates. The llama-embed-nemotron-8b model card says plainly that it's for non-commercial and research use only (model card). PolDense, stella-pl and the Polish rerankers are released under the Gemma licence, which comes with its own usage rules. Have a lawyer check the licence before you deploy.
Rerankers: where the biggest gain is
An embedding model encodes the question and the passage separately, so it has to be fast. A reranker (cross-encoder) reads the question and the passage together and judges whether the passage answers the question. It's slower, so it only scores the top few dozen candidates from the first stage.
In its reranker section, PIRB evaluates on 1,000 queries per dataset, with mmlw-retrieval-roberta-large as the default first stage. The results:
| Reranker | Parameters | PIRB average | Licence |
|---|---|---|---|
| BAAI/bge-reranker-v2.5-gemma2-lightweight | 9.2B | 66.68 | Gemma |
| Qwen/Qwen3-Reranker-8B | 8.2B | 65.77 | Apache 2.0 |
| sdadas/polish-reranker-roberta-v3 | 443M | 65.17 | Gemma |
| Qwen/Qwen3-Reranker-4B | 4.0B | 62.80 | Apache 2.0 |
| BAAI/bge-reranker-v2-m3 | 568M | 61.69 | Apache 2.0 |
| Qwen/Qwen3-Reranker-0.6B | 596M | 53.18 | Apache 2.0 |
For reference, the first stage on its own (mmlw-retrieval-roberta-large) scores 58.46 on the full data. The comparison is approximate, since rerankers were scored on a subset of queries, but the direction is clear:
- A good reranker adds several points. polish-reranker-roberta-v3 handles contexts up to 8,192 tokens, and its authors report that it beats Qwen3-Reranker-8B on 18 of 41 tasks with 5% of the parameters (model card).
- A weak reranker can hurt. Qwen3-Reranker-0.6B averages less in Polish than the first stage alone. Measure a reranker; don't bolt one on blindly.
- A better first stage helps the reranker. The same polish-reranker-roberta-v3 with BGE-Multilingual-Gemma2 as the first stage reaches 66.21 instead of 65.17.
The reranker score is also a good signal for the threshold below which the system answers “I don't know”. We cover that in our post on a knowledge base that says “I don't know”.
Hybrid search: don't throw out BM25
On PIRB, BM25 alone scores 41.85, but BM25 combined with mmlw-retrieval-roberta-large reaches 59.14, more than the same model alone (58.46). Full text catches what vectors struggle with: article numbers, case references, proper names, product codes.
The condition is that full-text search understands Polish inflection. In Elasticsearch that's the Stempel plugin, with the polish analyser and the polish_stem and polish_stop filters (Elastic docs). The two result lists are usually merged with RRF, which looks at ranks rather than scores (Cormack et al., 2009). BGE-M3 can return a dense vector and term weights in one pass (model card), but on PIRB its hybrid mode (55.75) doesn't improve on the dense version.
Practical criteria beyond the leaderboard
| Criterion | What to check | Example |
|---|---|---|
| Context length | How many tokens of a passage the model reads before cutting the rest | multilingual-e5-large: 512; BGE-M3 and PolDense: about 8k; Qwen3-Embedding: 32k |
| Prefixes and instructions | Without the right prefix, quality drops | PolDense: [query]: before the question; multilingual-e5: query: and passage: ; Qwen3-Embedding: a task instruction, typically worth 1 to 5% according to its authors |
| Vector size | Index memory grows linearly with dimensions | A million passages at 4,096 float32 values is about 16 GB; at 1,024 it's about 4 GB. Qwen3-Embedding lets you shorten the vector anywhere from 32 to 4,096 |
| Model size | Cost and time to index the whole base, hardware for queries | 8 billion parameters in BF16 is about 15 GB of weights alone; PolDense-17M to 68M is meant to run on CPU |
| Languages | Whether documents and questions mix Polish and English | PolDense's first training stage aligned Polish and English; EuroDense covers 9 languages |
| Serving | Whether the model runs in your inference server | The PolDense card gives a vLLM command (--runner pooling --convert embed) |
| Licence | Commercial use, usage rules | Apache 2.0 and MIT hold no surprises; Gemma comes with usage rules; Nemotron 8B is non-commercial only |
One thing that's easy to forget: switching embedding models means re-embedding the entire base. Vectors from different models aren't comparable. So make the choice carefully up front, and record the model version in the index metadata.
How to measure it on your own data
The leaderboard tells you which model is good on average across 41 tasks. Your knowledge base is one task: your documents and your questions. The test doesn't need to be big:
- Collect about a hundred real questions from the people who'll use the base. For each, note the passage that answers it.
- Add questions the documents can't answer. You'll need them to calibrate the reranker threshold.
- Index the same passages with two or three top models that fit your licence and size limits.
- Measure recall@10 and NDCG@10 for each model, for the first stage alone and with the reranker.
- Look at the misses by hand. Often the culprit turns out to be chunking or OCR, not the model. We write about that in our post on data preparation for AI.
Selection checklist
- Checked the current PIRB leaderboard, both the model and reranker sections.
- Picked 2-3 candidates whose licences fit your use case.
- Matched the model's context length to your passage size.
- Set prefixes and instructions as the model card says.
- Full-text search uses a Polish inflection analyser.
- List fusion (e.g. RRF) and a reranker over the top few dozen candidates.
- A test set: about a hundred questions with expected sources, plus unanswerable ones.
- Measured on your own data, not just read off the leaderboard.
- Model version recorded in index metadata, with a plan for re-embedding if you switch.
How we do it
In our knowledge bases, retrieval is hybrid: vectors plus a full-text index that handles Polish inflection, merged and ordered by a reranking model. The whole chain runs on a server inside the client's network, on NVIDIA Blackwell GPUs, so documents never go to external APIs. You can see it from scan to citation on our Knowledge bases page, and order a trial on your own documents through the demo.
We have no single favourite model. We pick the embedding model and the reranker for each project's needs: the kind of documents, the language, the volume and the hardware the system will run on, and we check the choice on a sample of the client's data.
Sources
- PIRB: Polish Information Retrieval Benchmark, leaderboard (Hugging Face, as of 30 Sep 2026)
- Dadas, Perełkiewicz, Poświata: PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods (arXiv, 2024)
- Dadas, Poświata, Grębowiec, Perełkiewicz: Parameter-Efficient Retrievers for Polish and European Languages (arXiv, 2026)
- OPI PIB: PolDense, a new generation of Polish information retrieval models
- Poświata, Dadas, Perełkiewicz: PL-MTEB: Polish Massive Text Embedding Benchmark (arXiv, 2024)
- OPI-PIB/PolDense-1B model card
- Qwen/Qwen3-Embedding-8B model card
- BAAI/bge-m3 model card
- intfloat/multilingual-e5-large model card
- sdadas/polish-reranker-roberta-v3 model card
- nvidia/llama-embed-nemotron-8b model card
- Elastic: Stempel Polish analysis plugin
- Cormack, Clarke, Büttcher: Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods (SIGIR 2009)
