Quick answer

Start with the PIRB leaderboard: 41 Polish retrieval tasks put together by a team at OPI PIB, Poland's National Information Processing Institute. As of 30 September 2026, PolDense-1B leads with 64.11 NDCG@10, ahead of models with 8-9 billion parameters. Popular multilingual models do worse in Polish: Qwen3-Embedding-8B scores 60.93, multilingual-e5-large 57.29 and BGE-M3 55.98. The biggest gain comes from reranking: the Polish polish-reranker-roberta-v3 (443M) reaches 65.17, close to Qwen3-Reranker-8B. Use the leaderboard to get down to two or three candidates, then let a test on about a hundred questions from your own documents decide.

Why model choice matters more in Polish

An embedding model turns a passage of text into a vector, and retrieval finds the passages whose vectors sit closest to the question's vector. If the model doesn't understand the language of your documents well, even the best language model at the end of the chain gets the wrong passages. We explain the whole chain in how RAG makes AI smarter.

Polish makes this harder in a few ways:

  • Inflection. “Umowa”, “umowy”, “umowie”, “umową” are all forms of the word for “contract”. Full-text search without a Polish analyser treats them as different words.
  • Less training data. Multilingual models learn mostly from English text. Polish gets a sliver of the data, and specialist vocabulary (law, medicine, public administration) gets even less.
  • Mixed documents. Polish companies mix Polish and English: bilingual contracts, technical documentation, correspondence. The model has to find a Polish passage for an English question and the other way round.

English benchmarks (such as English MTEB) say little about any of this, which is why Polish leaderboards are worth a look.

What PIRB is and how to read it

PIRB (Polish Information Retrieval Benchmark) was published in 2024 by Sławomir Dadas, Michał Perełkiewicz and Rafał Poświata of OPI PIB (Dadas et al., 2024). It covers 41 retrieval tasks in five groups:

  • MaupQA: an automatically created Polish question answering dataset (from IPI PAN);
  • PolEval-2022: tasks from the PolEval challenge, including legal questions;
  • BEIR-PL: BEIR datasets machine-translated into Polish (by Wrocław University of Science and Technology);
  • Web Datasets: real questions and answers from Polish websites, including legal and medical ones;
  • Other: including MFAQ and GPT-exams.

The main metric is NDCG@10, which scores whether the right passages appear in the top ten results and how high. The leaderboard has three sections: standalone models, hybrid retrieval, and two-stage retrieval with a reranker.

For a broader view there's PL-MTEB from the same authors: 30 tasks in five categories (classification, clustering, pair classification, retrieval and semantic similarity), merged into MTEB (Poświata et al., 2024). For RAG, retrieval is what matters most, and PIRB covers it in the most depth.

One note for completeness: PIRB, PolDense, the mmlw family and the Polish rerankers all come from the same team. The benchmark is open and the evaluation code is public (github.com/sdadas/pirb), but it's one more reason to confirm leaderboard results on your own data.

Embedding model leaderboard (as of 30 September 2026)

Average NDCG@10 across 41 tasks, computed from the PIRB leaderboard data. The “legal” column averages four legal tasks: two from PolEval-2022 (legal questions) plus eprawnik and specprawnik.

ModelParametersPIRB averagePIRB legalLicence
OPI-PIB/PolDense-1B1.0B64.1173.05Gemma
nvidia/llama-embed-nemotron-8b7.5B63.7369.39non-commercial only
BAAI/bge-multilingual-gemma29.2B63.2671.82Gemma
OPI-PIB/PolDense-400M395M63.2171.91Gemma
sdadas/stella-pl-retrieval-8k1.5B62.6968.38Gemma
OPI-PIB/PolDense-150M149M61.0768.54Gemma
Qwen/Qwen3-Embedding-8B7.6B60.9364.49Apache 2.0
Qwen/Qwen3-Embedding-4B4.0B59.1962.38Apache 2.0
sdadas/mmlw-retrieval-roberta-large435M58.4660.13Apache 2.0
intfloat/multilingual-e5-large560M57.2957.89MIT
BAAI/bge-m3 (dense)568M55.9859.80MIT
Qwen/Qwen3-Embedding-0.6B596M50.7150.80Apache 2.0
BM25 (no model)n/a41.8545.24n/a

What this tells you:

  • Size isn't everything. PolDense-400M scores close to models about 20 times its size. According to its model card, PolDense-150M comes close to Qwen3-Embedding-8B (model card), and OPI PIB says the smallest variants (17M, 32M, 68M) are designed to run on CPUs and edge devices (OPI PIB).
  • Multilingual models lose ground in Polish. Qwen3-Embedding-8B ranked first on multilingual MTEB in June 2025 (70.58) (model card). In Polish it's solid, but behind models trained for Polish. The small 0.6B version does worse than multilingual-e5-large.
  • Domain widens the gaps. On specprawnik (questions from a Polish legal website), PolDense-1B scores 46.6, Qwen3-Embedding-8B 31.3 and BGE-M3 25.5. If your documents are specialised, the overall average may understate the difference.
  • Licences rule out candidates. The llama-embed-nemotron-8b model card says plainly that it's for non-commercial and research use only (model card). PolDense, stella-pl and the Polish rerankers are released under the Gemma licence, which comes with its own usage rules. Have a lawyer check the licence before you deploy.

Rerankers: where the biggest gain is

An embedding model encodes the question and the passage separately, so it has to be fast. A reranker (cross-encoder) reads the question and the passage together and judges whether the passage answers the question. It's slower, so it only scores the top few dozen candidates from the first stage.

In its reranker section, PIRB evaluates on 1,000 queries per dataset, with mmlw-retrieval-roberta-large as the default first stage. The results:

RerankerParametersPIRB averageLicence
BAAI/bge-reranker-v2.5-gemma2-lightweight9.2B66.68Gemma
Qwen/Qwen3-Reranker-8B8.2B65.77Apache 2.0
sdadas/polish-reranker-roberta-v3443M65.17Gemma
Qwen/Qwen3-Reranker-4B4.0B62.80Apache 2.0
BAAI/bge-reranker-v2-m3568M61.69Apache 2.0
Qwen/Qwen3-Reranker-0.6B596M53.18Apache 2.0

For reference, the first stage on its own (mmlw-retrieval-roberta-large) scores 58.46 on the full data. The comparison is approximate, since rerankers were scored on a subset of queries, but the direction is clear:

  • A good reranker adds several points. polish-reranker-roberta-v3 handles contexts up to 8,192 tokens, and its authors report that it beats Qwen3-Reranker-8B on 18 of 41 tasks with 5% of the parameters (model card).
  • A weak reranker can hurt. Qwen3-Reranker-0.6B averages less in Polish than the first stage alone. Measure a reranker; don't bolt one on blindly.
  • A better first stage helps the reranker. The same polish-reranker-roberta-v3 with BGE-Multilingual-Gemma2 as the first stage reaches 66.21 instead of 65.17.

The reranker score is also a good signal for the threshold below which the system answers “I don't know”. We cover that in our post on a knowledge base that says “I don't know”.

Hybrid search: don't throw out BM25

On PIRB, BM25 alone scores 41.85, but BM25 combined with mmlw-retrieval-roberta-large reaches 59.14, more than the same model alone (58.46). Full text catches what vectors struggle with: article numbers, case references, proper names, product codes.

The condition is that full-text search understands Polish inflection. In Elasticsearch that's the Stempel plugin, with the polish analyser and the polish_stem and polish_stop filters (Elastic docs). The two result lists are usually merged with RRF, which looks at ranks rather than scores (Cormack et al., 2009). BGE-M3 can return a dense vector and term weights in one pass (model card), but on PIRB its hybrid mode (55.75) doesn't improve on the dense version.

Practical criteria beyond the leaderboard

CriterionWhat to checkExample
Context lengthHow many tokens of a passage the model reads before cutting the restmultilingual-e5-large: 512; BGE-M3 and PolDense: about 8k; Qwen3-Embedding: 32k
Prefixes and instructionsWithout the right prefix, quality dropsPolDense: [query]: before the question; multilingual-e5: query: and passage: ; Qwen3-Embedding: a task instruction, typically worth 1 to 5% according to its authors
Vector sizeIndex memory grows linearly with dimensionsA million passages at 4,096 float32 values is about 16 GB; at 1,024 it's about 4 GB. Qwen3-Embedding lets you shorten the vector anywhere from 32 to 4,096
Model sizeCost and time to index the whole base, hardware for queries8 billion parameters in BF16 is about 15 GB of weights alone; PolDense-17M to 68M is meant to run on CPU
LanguagesWhether documents and questions mix Polish and EnglishPolDense's first training stage aligned Polish and English; EuroDense covers 9 languages
ServingWhether the model runs in your inference serverThe PolDense card gives a vLLM command (--runner pooling --convert embed)
LicenceCommercial use, usage rulesApache 2.0 and MIT hold no surprises; Gemma comes with usage rules; Nemotron 8B is non-commercial only

One thing that's easy to forget: switching embedding models means re-embedding the entire base. Vectors from different models aren't comparable. So make the choice carefully up front, and record the model version in the index metadata.

How to measure it on your own data

The leaderboard tells you which model is good on average across 41 tasks. Your knowledge base is one task: your documents and your questions. The test doesn't need to be big:

  1. Collect about a hundred real questions from the people who'll use the base. For each, note the passage that answers it.
  2. Add questions the documents can't answer. You'll need them to calibrate the reranker threshold.
  3. Index the same passages with two or three top models that fit your licence and size limits.
  4. Measure recall@10 and NDCG@10 for each model, for the first stage alone and with the reranker.
  5. Look at the misses by hand. Often the culprit turns out to be chunking or OCR, not the model. We write about that in our post on data preparation for AI.

Selection checklist

  • Checked the current PIRB leaderboard, both the model and reranker sections.
  • Picked 2-3 candidates whose licences fit your use case.
  • Matched the model's context length to your passage size.
  • Set prefixes and instructions as the model card says.
  • Full-text search uses a Polish inflection analyser.
  • List fusion (e.g. RRF) and a reranker over the top few dozen candidates.
  • A test set: about a hundred questions with expected sources, plus unanswerable ones.
  • Measured on your own data, not just read off the leaderboard.
  • Model version recorded in index metadata, with a plan for re-embedding if you switch.

How we do it

In our knowledge bases, retrieval is hybrid: vectors plus a full-text index that handles Polish inflection, merged and ordered by a reranking model. The whole chain runs on a server inside the client's network, on NVIDIA Blackwell GPUs, so documents never go to external APIs. You can see it from scan to citation on our Knowledge bases page, and order a trial on your own documents through the demo.

We have no single favourite model. We pick the embedding model and the reranker for each project's needs: the kind of documents, the language, the volume and the hardware the system will run on, and we check the choice on a sample of the client's data.

Sources