Quick answer
Choose a model by task, hardware and license, not by leaderboard. Shortlist two or three candidates and test them on 50-100 of your own examples, especially in languages other than English, where public tests contradict each other. In late August 2026 a sensible starting point for RAG and agents on one GPU is a dense 27-31B model under Apache 2.0 (for example Qwen3.8-27B or Gemma 4 31B). Polish-focused models (Bielik v3, PLLuM) are worth adding to the test when Polish matters and hardware is smaller. Models above 100 billion parameters (Mistral Small 4, GLM-5.3-Flash) need more memory but bring more quality on hard tasks.
We checked every model in this article against its Hugging Face model card on 31 August 2026. This field changes every quarter, so we update the article.
Which open model families should you know?
| Family | Example models (Hugging Face) | Architecture and context | License | Notes |
|---|---|---|---|---|
| Qwen (Alibaba) | Qwen3.8-27B, Qwen3.6-35B-A3B, Qwen3.8-Flash-Next | 27B dense with image and video understanding, 262k context (up to 1M); 35B MoE; Flash-Next: 125B with 6B active plus 51B n-gram embeddings | Apache 2.0; Flash-Next: Qwen Community License 1.0 | reasoning mode on by default, can be switched off per request; official FP8 versions |
| Gemma 4 (Google DeepMind) | Gemma 4 31B, 26B A4B, 12B, E4B, E2B | dense and MoE, 128k context (small) or 256k; images on all, audio on E2B, E4B and 12B | Apache 2.0 | 35+ languages out of the box, pre-trained on 140+; official 4-bit QAT versions |
| GLM (Z.ai) | GLM-4.7-Flash, GLM-5.3-Flash, GLM-5.3 | 4.7-Flash: 30B-A3B MoE; 5.3-Flash: 320B with 18B active, multimodal; 5.3: about 750B | MIT; GLM-5.3: own license | strong at code and agent tasks according to the vendor; the card lists English and Chinese |
| Mistral (Mistral AI) | Ministral 3 14B, Mistral Small 4, Mistral Medium 3.5 | Ministral: 3B, 8B, 14B; Small 4: 119B MoE with 6.5B active, 256k context; Medium 3.5: 128B dense | Apache 2.0; Medium 3.5: modified MIT | Small 4 combines instruct, reasoning and coding modes; official NVFP4 version |
| Bielik (SpeakLeash, ACK Cyfronet AGH) | Bielik-11B-v3.0-Instruct, Bielik-PL-11B-v3.0-Instruct, Bielik-Minitron-7B-v3.0 | 11B dense (Minitron 7B) | Apache 2.0, access after accepting conditions | trained with an emphasis on Polish and 32 European languages; the PL version has a tokenizer optimized for Polish |
| PLLuM (PLLuM consortium, HIVE AI) | PLLuM-12B-instruct-2512, Llama-PLLuM-70B-instruct-2512 | 12B based on Mistral Nemo; 70B based on Llama 3.1 | 12B: Apache 2.0; 70B: Llama 3.1 license | Polish data and hand-written instructions, geared towards public administration |
| gpt-oss (OpenAI) | gpt-oss-20b, gpt-oss-120b | MoE, 4-bit weights (MXFP4) | Apache 2.0 | a good reference point in tests |
This table isn't a ranking. It's a list of candidates from which you pick two or three for your own test.
Why isn't a leaderboard enough?
Vendor benchmarks are run on the vendor's own settings, usually in English or Chinese. In Polish, the language we work in most, the picture is mixed:
- The Oxido test from March 2026 compared 12 models on 20 tasks in 10 categories (emails, business advice, language correctness, history and culture). The Polish models, Bielik and PLLuM, ended up near the bottom (Bankier.pl). The result sparked a debate, because the test was small and mostly scored general tasks.
- The PLCC benchmark from Poland's National Information Processing Institute has 600 hand-crafted questions on history, geography, culture, art, grammar and vocabulary. The authors tested more than 30 models and found that commercial models hold a clear lead, but that small Bielik can match much larger multilingual models on understanding Polish culture (Dadas et al., 2025).
- The Open PL LLM Leaderboard from SpeakLeash aggregates results from many Polish tasks (Hugging Face). It's a good first filter, but its tasks aren't your tasks.
The takeaway applies to any language: a leaderboard tells you which models are worth testing. It doesn't tell you which one will be best in your knowledge base, with your documents and your tools.
License: what to check before deploying
"Open model" doesn't always mean "use it however you like". Here are the differences we found in the model cards:
| License | Examples | Watch out for |
|---|---|---|
| Apache 2.0 | Qwen3.8-27B, Qwen3.6, Gemma 4, Mistral Small 4, Ministral 3, Bielik v3, PLLuM-12B, gpt-oss | no extra conditions; Bielik asks you to accept access conditions on Hugging Face |
| MIT | GLM-4.7-Flash, GLM-5.3-Flash | no extra conditions |
| GLM-5.3 license | GLM-5.3 | a Model-as-a-Service provider with revenue above $10 billion a year must pass a Z.ai security review |
| Qwen Community License 1.0 | Qwen3.8-Flash-Next | a Model-as-a-Service or AI work assistant business needs a separate license from Qwen; internal use is allowed |
| Modified MIT | Mistral Medium 3.5 | no rights for companies with monthly revenue above $20 million |
| Llama 3.1 | Llama-PLLuM-70B | Meta's license terms |
If you're building a product in which customers use the model, the license matters from day one. This is not legal advice: if in doubt, have a lawyer read the license.
Which model will fit your hardware?
Memory is driven by all parameters, speed by the active ones. A 35B-A3B MoE model generates as fast as a 3B model but takes as much memory as a 35B one. Detailed weight and KV cache calculations are in our guide how much VRAM for a local LLM, part 1 and part 2. In short:
| Hardware | Sensible candidates |
|---|---|
| Laptop, 16 GB card | Gemma 4 E4B or 12B, Ministral 3 8B, Bielik-Minitron-7B, gpt-oss-20b |
| 24-32 GB card | Gemma 4 12B, Gemma 4 26B A4B in 4-bit, Bielik 11B, PLLuM-12B, Qwen3.6-35B-A3B in 4-bit (on 32 GB) |
| 96 GB card | Qwen3.8-27B or Gemma 4 31B in FP8 for many users, GLM-4.7-Flash, Mistral Small 4 in NVFP4 |
| Several cards or B200/B300-class GPUs | GLM-5.3-Flash, Mistral Medium 3.5, larger MoE models |
The weight format matters too: official FP8 (Qwen), NVFP4 (Mistral) or 4-bit QAT (Gemma) releases usually hold quality better than community quantizations. More on that in FP8, NVFP4 or AWQ.
How to test models on your own examples
This is the most important part of the choice, and it usually takes a few days.
- Collect 50-100 real requests. From ticket history, from email, from pilot users. Add hard cases: questions the documents can't answer, vague ones, ones with typos, ones that mix languages.
- Write down the expected answers or the criteria for a correct answer. Without them, scoring comes down to gut feeling.
- Define criteria. Factual correctness, faithfulness to sources, language quality, format (JSON, tool calls), length, response time.
- Pick two or three candidates from the table above that fit your hardware and license.
- Run them all with identical settings: the same system prompt, the same RAG excerpts, the same temperature, the same engine (vLLM or SGLang).
- Score the results. By hand for the full set, or with a judge model whose ratings you check by hand on a sample of 20-30 answers.
- Keep the set as a regression test. Run it for every new version of the model, the engine or the prompt.
What to look at beyond correctness:
- Reasoning mode. Some models "think" before answering by default. That helps on hard tasks but adds latency and burns tokens. Check whether you can switch it off per request.
- Language of reasoning versus answer. A model can reason in English and answer in Polish or German. Under a token limit, the answer may not fit.
- Tool calls. For agents, what matters is whether the model calls the right tool with correct arguments, not just how good the text reads.
- Sticking to sources. In a knowledge base the model should answer from the documents and say "I don't know" when the answer isn't there.
Which model for which job?
| Task | What to look at | Candidates to test |
|---|---|---|
| RAG, knowledge base | faithfulness to sources, long context, language quality | Qwen3.8-27B, Gemma 4 31B, Mistral Small 4, Bielik 11B v3 for Polish |
| Agents with tools | quality of tool calls, long sessions | Qwen3.8-27B, GLM-4.7-Flash or GLM-5.3-Flash, Mistral Small 4 |
| Classification, extraction | stable format, speed, fine-tunability | small models (Gemma 4 12B, Ministral 3, Bielik-Minitron) with LoRA |
| Scanned documents and images | image understanding | Qwen3.8-27B, Gemma 4, GLM-5.3-Flash |
| Edge devices, laptops | size, audio | Gemma 4 E2B or E4B |
For classification and extraction, a small model fine-tuned with LoRA often beats a big model with a prompt. We walk through an example in small model with LoRA or big model with a prompt.
How we do it
We serve models from the Qwen, Gemma, GLM and Mistral families on our own NVIDIA Blackwell GPUs and compare successive generations in practice. Every model change is preceded by a test on a set of real agent tasks and knowledge-base questions. A few observations from those tests:
- A better score on the vendor's benchmark didn't always mean a better agent. We've seen a model with reasoning on by default spend its token budget on long English reasoning and run out before giving the Polish answer.
- The same model in a different quantization, or on a different engine version, could behave differently on tool calls while ordinary answers looked identical.
- In one of our comparisons, a dense model of around 30B stuck to the system prompt and wrote better Polish than an MoE model with a few billion active parameters, and answered faster because the other model reasoned by default. That's one pair of models on one task, not a rule.
If you'd like to compare a few models on your own data, you can request a demo. How we select and fine-tune models is described on our models and fine-tuning page, and when to choose an open model over an API at all is covered in how much does an on-prem LLM cost.
Model selection checklist
- The task fits in one sentence and has a measurable success criterion.
- You know your GPU memory budget and the number of concurrent users.
- The candidates' licenses fit how you'll use the model (internally or in a customer-facing product).
- You have 50-100 examples with expected answers.
- You test two or three models with identical settings.
- You check language quality, faithfulness to sources and tool calls separately.
- You measure response time at realistic user numbers.
- The test set stays on as a regression test for future versions.
Sources
Model cards on Hugging Face (as of 31 August 2026):
- Qwen/Qwen3.8-27B, Qwen/Qwen3.8-27B-FP8, Qwen/Qwen3.6-35B-A3B, Qwen/Qwen3.8-Flash-Next
- google/gemma-4-31B-it, google/gemma-4-31B-it-qat-w4a16-ct
- zai-org/GLM-4.7-Flash, zai-org/GLM-5.3-Flash, zai-org/GLM-5.3
- mistralai/Ministral-3-14B-Instruct-2512, mistralai/Mistral-Small-4-119B-2603, mistralai/Mistral-Medium-3.5-128B
- speakleash/Bielik-11B-v3.0-Instruct, speakleash/Bielik-PL-11B-v3.0-Instruct
- CYFRAGOVPL/PLLuM-12B-instruct-2512
- openai/gpt-oss-20b, openai/gpt-oss-120b
Polish-language tests and benchmarks:
- Bankier.pl: Polish AI lost to the giants; Bielik and PLLuM scored poorly in tests (17 Mar 2026, in Polish)
- Dadas et al.: Evaluating Polish linguistic and cultural competency in large language models (2025)
- SpeakLeash: Open PL LLM Leaderboard
- Bielik 11B v3: Multilingual Large Language Model for European Languages
