Quick answer
RAG (Retrieval-Augmented Generation) is a technique where a language model, before answering, searches current sources such as your company's documents and databases and answers based on what it finds, often with a citation. This lets AI use up-to-date and private data without retraining the model, and it makes things up less often. Less often isn't never: RAG reduces hallucinations but doesn't eliminate them, so good systems show their sources and can say "I don't know."
Why does a language model without RAG guess?
Because it only knows what it saw during training. A large language model (LLM) learns to understand and generate language by analyzing vast amounts of text: documents, articles, books and conversations. It learns patterns of words, sentences and context, which is why it sounds natural on almost any topic.
That training is very expensive. According to the Stanford AI Index, the estimated compute cost of training GPT-4 was about $78 million, and Gemini Ultra about $191 million (Stanford HAI, AI Index 2024). Nobody is going to repeat that every time your price list or a procedure changes.
The result: a model's knowledge is frozen at a point in time and doesn't include data it has never seen, such as your contracts, procedures or inventory. It's like a student taking an exam from memory alone: they can only answer from what they studied weeks ago, with no way to check the facts. A model trained in early 2024 doesn't know what happened in late 2025. And when it doesn't know something, it often answers anyway, confidently. That's a hallucination: an answer that sounds credible but is made up.
How does RAG change the game?
RAG turns a closed-book exam into an open-book one. Instead of relying only on what it learned in training, the model first checks current, verified sources and builds its answer on them. The concept and first architecture were described in 2020 by a Facebook AI Research team, combining knowledge stored in the model with an external document index searched by a separate retrieval model (Lewis et al., 2020).
Key benefits:
- Fewer hallucinations. The model bases its answer on retrieved content instead of guessing from hazy training memories.
- Transparency. The answer can point to the exact document and passage it came from. Information can be checked, which builds trust.
- The right to say "I don't know." If the sources don't contain the answer, the system can admit it. "I don't have that information" is far better than a confidently wrong answer.
- Fresh knowledge without training. Change a document in the base, and from the next question on, answers reflect the change.
How does RAG work step by step?
RAG works in four steps: one preparation step and three that run for every question.
Step 0: Indexing, or preparing the knowledge. Documents are split into chunks, and each chunk is turned into an embedding: a vector of numbers that represents its meaning. The vectors go into a vector database that can quickly find chunks with similar meaning. This step determines everything that follows: a badly read scan or a table cut in half becomes a chunk the model can't understand. That's why RAG often starts with OCR, which we cover in what modern OCR is.
Step 1: Retrieval, or finding what's relevant. The system searches the sources for the information most related to the question. Sources can be public (websites, articles, public registries) or internal (knowledge bases, documents, product catalogs, business systems). Example: for "When will my package arrive?" the system checks the status of that specific shipment in the carrier's database.
Step 2: Augmentation, or adding context. The retrieved data is attached to the question. The model now has both the question and facts from real sources. Example: the system finds that the package is in transit from Chicago, was last scanned 2 hours ago and is scheduled for delivery on October 22, and attaches this to the question.
Step 3: Generation, or writing the answer. Only now does the model write the answer, combining the question with the retrieved data, and it can cite the source. Example: "Your package is in transit from Chicago and will arrive tomorrow, October 22, by 8 p.m. [Source: shipment database, updated 2 hours ago]."
Does RAG eliminate hallucinations?
No. RAG reduces them significantly but doesn't remove them. A Stanford study shows this well: researchers tested commercial legal research tools marketed as hallucination-free thanks to RAG. The tools were wrong less often than general-purpose GPT-4, but they still hallucinated in 17-33% of cases (Magesh et al., Journal of Empirical Legal Studies, 2025).
Where do errors come from? Retrieval can return the wrong passage, miss the right one or return a passage without its context, and the model can combine correct facts incorrectly. What helps in practice:
- Better retrieval: combining meaning-based (vector) search with keyword search, plus reranking, where a separate model reorders the results.
- A citation for every claim, so users can check the source in one click.
- Grounding checks: automatically verifying that the answer follows from the retrieved passages before it reaches the user.
- Permission to say "I don't know," written into the model's instructions and verified in testing.
- A test set of questions with expected answers, run after every change.
RAG, fine-tuning or long context: which should you choose?
RAG adds knowledge to a model, fine-tuning changes its behavior. This is a common misunderstanding: companies want to "teach the model our documents," but what they usually need is RAG.
| Criterion | RAG | Fine-tuning | Long context (pasting documents into the prompt) |
|---|---|---|---|
| Purpose | Answers based on current documents | Changing style or format, specializing in a task | Working with a few known documents at once |
| Updating knowledge | Change the document in the base | Retrain | Paste the new version |
| Sources and citations | Yes, naturally | No | Partly |
| Knowledge scale | Very large (millions of chunks) | Limited | Limited by the context window and cost |
| Access control | Per document or chunk | None, knowledge is in the weights | Manual |
| Typical example | Knowledge base, helpdesk, legal assistant | A model that writes in your company's style | Analyzing a single contract |
These approaches can be combined. We show how we fine-tune open models on our Models & fine-tuning page.
How do we build RAG in practice?
We deploy our agentic knowledge base in finance, insurance and pharma. A few decisions from this project show where quality in RAG really comes from:
- Documents: PDFs, Excel spreadsheets and scans go into one pipeline. OCR reads scans in Polish and English, tables are converted to text with column headers, charts are described by a vision model, and each chunk gets a prefix with the document and section name.
- Language-aware search: vector search runs in parallel with Polish full-text search that reduces words to their stems, so inflected forms of the same word match. A reranker orders the results, and the two lists are merged using RRF (reciprocal rank fusion).
- An agent instead of a single search: the model decides whether it has enough context and, if needed, pulls in a whole chapter or checks an exact phrase. This approach is called agentic RAG, and we describe it in agentic RAG: from answering to acting.
- Answers you can verify: every claim has a citation to a passage, and before sending, the answer passes a grounding check and a hallucination detector.
- Scale: the knowledge base built on a public drug registry contains over 2 million chunks from 22,786 documents.
How does RAG protect confidential data?
Only if you design it to. RAG reaches into company documents, so retrieval must respect permissions: a sales rep shouldn't get an excerpt from a board-level contract just because it matched the question. Permission filters are applied at retrieval time, before a passage ever reaches the model. We cover permissions and data protection in RAG systems in secure RAG: permissions and data access.
The second decision is where processing happens. Our knowledge base runs on-premise, on our own GPUs with local models, so documents and questions never leave the company. We show what AI on your own infrastructure looks like on our AI infrastructure page.
Where does RAG make the most sense?
Anywhere answers have to be found in documents or data that change. RAG is why a bank's assistant can answer questions about your account, a store's assistant knows what's in stock right now, and a company AI can search procedures and policies. In the same way, a medical assistant can pull current guidelines, a legal tool can cite case law, and a helpdesk can check an order's status.
Where to start:
- Pick one knowledge area with a clear owner, such as customer service procedures or product documentation.
- Write down 30-50 real questions people ask today, together with the correct answers. That's your test set.
- Clean up the documents: remove outdated versions and make sure scans are legible and structured.
- Settle permissions and where data is processed before the first prototype, not after.
- Measure quality on the test set, and only then roll it out more widely.
Without RAG, a chatbot gives everyone the same generic answers from outdated knowledge. With RAG, it becomes an assistant that pulls current information tailored to you and your situation, and shows where it came from. RAG turns guessing into knowing, and that makes all the difference.
Sources
- Lewis P. et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv 2020
- Magesh V. et al.: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies, 2025 (arXiv preprint)
- Stanford HAI: AI Index: State of AI in 13 Charts (2024)