Quick answer

A knowledge base that doesn't make things up needs three safeguards. First, the model writes only from the passages it was given and cites a source for every sentence, and each citation leads to a specific sentence in the document, not to the paragraph next to it. Second, retrieval has a relevance threshold: when no passage clears it after reranking, the system says “I don't know” instead of building an answer from whatever came back. Third, the permission filter works inside the index, so the model never sees documents the person asking isn't allowed to read. The threshold has to be calibrated on your own questions, including ones the documents can't answer.

Why does RAG make things up when it has the documents?

RAG (Retrieval-Augmented Generation) is a technique where the model searches a document base before answering and bases its answer on what it finds. We explain the basic version in how RAG makes AI smarter. This post is about what happens when retrieval brings back too little, or the wrong thing.

Having documents in the context is not enough to keep a model honest. The research is fairly consistent:

  • Citations that don't support the claim. A Stanford team audited answers from popular generative search engines. Only 51.5% of generated sentences were fully supported by citations, and only 74.5% of citations supported the sentence they were attached to (Liu, Zhang, Liang, 2023). One citation in four looked credible but didn't back up the text beside it.
  • Incomplete support even at the top. On the ALCE benchmark, the best models lacked complete citation support half of the time on the ELI5 dataset (Gao et al., 2023).
  • Answering instead of abstaining. Google researchers showed that large models answer well when the context is sufficient, but when it isn't, they often give a wrong answer rather than hold back (Joren et al., 2024).
  • Specialist tools get it wrong too. Commercial legal research tools built on RAG hallucinated between 17% and 33% of the time: less than general-purpose GPT-4, but far from zero (Magesh et al., 2025).

The takeaway is simple. Giving the model documents isn't enough. You need mechanisms that force it to write from sources, let people check that it did, and give the system permission to refuse.

Safeguard one: a citation for every sentence

In our knowledge base, the model writes only from the passages in its context and cites a source for every sentence. Clicking a citation opens the document and highlights the sentence it came from. You can see this in the story on our Knowledge bases page, which uses a fictional law firm as the example.

Why a citation per sentence and not a list of sources under the answer?

  • It can be checked. “Sources: contract.pdf, terms.pdf” doesn't say which sentence came from where. A per-sentence citation lets you verify each claim on its own, in seconds.
  • It keeps the model disciplined. When every sentence has to be pinned to a passage, it's harder for the model to slip in a sentence of its own. A sentence with no citation stands out immediately.
  • It points to a place, not a file. The citation leads to a specific sentence in the document. That means document processing has to keep each passage's position (page, paragraph, character offset), not just its text.

A citation is only the model's claim. Whether the cited passage really supports the sentence has to be checked separately. This step is called a groundedness (or attribution) check. The literature does it in two ways: a natural language inference (NLI) model that judges whether the passage entails the sentence, or a second language model pass acting as a judge (Rashkin et al., 2021; the faithfulness metric in RAGAS). A sentence that fails the check can be removed, flagged, or sent back for regeneration.

For us this step is not a single filter but a set of gates. A precision gate checks that the sentence is actually supported by the cited passage, and a hallucination gate catches claims that appear in no source at all. A sentence that fails the gates does not reach the answer as a fact. A human always stays in the loop: answers on sensitive matters are approved by someone on the client's team. We also collect trajectories, a record of what the system retrieved, what it rejected and what it answered. We group them into datasets and grade them two ways: with a model acting as a judge (LLM as a judge) and by hand. That way the gate thresholds are set from graded cases, not by feel.

Safeguard two: a relevance threshold and “I don't know”

Retrieval always returns something. Even when the answer isn't in the base, a vector search will hand back the five “nearest” passages, because something is always nearest. If the model gets them with no warning, it does what models do: it writes a plausible answer from them.

So in our pipeline, retrieval works like this. The question is classified first, then we search the vectors and the full text in parallel, merge the two lists, and a reranking model reads each question and passage pair and sets the final order. When nothing clears the threshold, the answer is “I don't know”.

Why put the threshold on the reranker and not on vector similarity?

SignalWhat it measuresGood for a threshold?
Vector similarity (cosine)How close the question and passage sit in meaning space, each encoded separatelyNot really. Values depend on the model and text length, and “on the same topic” is not the same as “answers the question”
BM25 scoreKeyword overlapNo. The scale depends on the query and the corpus
Position after list fusion (RRF)Order, not strength of the matchNo. RRF deliberately looks only at ranks (Cormack et al., 2009)
Reranker score (cross-encoder)Whether this passage answers this question, judged on the pair read togetherYes, once calibrated on your own data

A reranker reads the question and the passage together, so it judges whether the passage answers the question, not just whether it's about the same thing. Popular rerankers produce a score you can map to the 0 to 1 range. BGE-reranker-v2-m3, for example, runs its score through a sigmoid (model card), and Qwen3-Reranker computes the probability of a “yes” answer (model card). That still isn't a calibrated probability that the answer is correct, but it's a good signal to threshold on.

How to pick the threshold

Pick it from data, not by feel:

  1. Collect questions the documents can answer, each with the expected source.
  2. Collect questions the documents can't answer: on the same topics, but with no answer in the base. Without them you can't measure whether the system knows how to refuse.
  3. Run both sets through retrieval and record the top reranker score for each question.
  4. Choose the threshold that balances two errors: false “I don't know” on answerable questions, and unsupported answers on unanswerable ones.
ThresholdWhat happensWhat users notice
Too highThe system refuses even though the answer is in the documents“This thing knows nothing”, and people go back to searching by hand
Too lowThe system answers from passages that only resemble the topicConfident, wrong answers with citations that don't fit
Tuned on a test setRefuses where there's no source, answers where there is one“I don't know” is rare and usually justified

A high rate of false refusals is rarely a threshold problem. It usually means something earlier is failing: OCR, chunking, or full-text search that doesn't handle Polish inflection. We cover document processing in our post on modern OCR.

What a good “I don't know” looks like

A bare “I don't know” is frustrating. A good refusal says more:

  • that no passage in the available documents answers the question;
  • which documents came closest, in case the user wants to look for themselves;
  • how to rephrase the question, or who to ask.

It also must not reveal that the answer exists in documents the person isn't allowed to see. Which brings us to the third safeguard.

Safeguard three: a department filter inside the index

Every passage in our knowledge base carries the department it belongs to in its metadata. The filter works inside the index itself, before anything is searched, so another department's files never reach the ranking or the model's context. Every question goes into the audit log: who asked, when, and which sources the answer relied on.

Why filter before retrieval and not after?

  • The model never sees what it shouldn't. If you only filter the results, a restricted document may already have influenced the ranking or landed in the context, and the model can leak something from it into the answer.
  • The threshold stays honest. When the only matching passage belongs to another department, the user gets “I don't know”, exactly as if the document didn't exist. They don't learn that an answer is out there somewhere.
  • The audit log means something. It shows which sources the answer relied on, and those sources were definitely available to the person asking.

We go deeper into permissions in our post on secure RAG data access.

What each mechanism solves, and what it doesn't

MechanismProtects againstWon't fix
Model writes only from passagesAnswers from the model's general, often outdated knowledgePoorly chosen passages
Citation for every sentenceClaims nobody can checkCitations that point to the wrong text
Groundedness checkCitations that don't support the sentenceErrors in the source documents themselves
Relevance threshold and “I don't know”Answers assembled from loosely related passagesWeak retrieval (false refusals go up)
Department filter in the indexLeaks between departments through ranking and contextWrong permissions set in the source system
Audit logNo record of who asked what and where the answer came fromAnything at all, if nobody reads it

None of these is enough on its own. Together they mean an answer either has a source you can check, or there is no answer.

Checklist: a knowledge base that doesn't make things up

  • The prompt tells the model to write only from the passages provided and to refuse when they're missing.
  • Every sentence in the answer cites a specific passage.
  • Each citation opens the document at the place the sentence came from (position kept during processing).
  • A groundedness check runs after generation: does the passage support the sentence?
  • Retrieval is hybrid (vectors plus full text with Polish inflection), with reranking.
  • The relevance threshold sits on the reranker score and is calibrated on a test set.
  • The test set includes questions the documents can't answer.
  • The permission filter works inside the index, before retrieval.
  • A refusal doesn't reveal documents the user can't access.
  • Every question, answer and source goes into the audit log.
  • After any change to the model, the threshold or document processing, the test set runs again.

How to check this on your own documents

The fastest way is a sample. We take a few hundred documents from one department and a list of questions people really ask, each with the source they expect. We build the base, and the result is a report: where the answer hits the source, where the system says “I don't know”, and why. The whole chain (OCR, indexes, embedding model, reranker and language model) runs on a server inside your network, on NVIDIA Blackwell GPUs.

If questions need several searches or pull from many documents, the next step is agentic RAG, where an agent decides for itself whether it has found enough. The rule stays the same: no source, no answer. The full chain from scan to citation is on our Knowledge bases page, and you can order a trial base on your own documents through the demo.

Sources