Quick answer

Agentic RAG is RAG (Retrieval-Augmented Generation) with an AI agent between the question and the answer: it decides where and how to search, evaluates the passages it finds, rewrites queries and pulls in missing context before generating a response. A 2025 research survey describes it as embedding autonomous agents into the RAG pipeline using reflection, planning, tool use and multi-agent collaboration (Singh et al., 2025). The difference from classic RAG: instead of one search in a fixed sequence, there is a loop that runs until the context is sufficient.

What is classic RAG and where does it fall short?

RAG is a technique in which a language model searches a document store before answering and grounds its answer in what it finds, rather than relying only on what it learned in training. The term was introduced in 2020 by a Facebook AI Research team with researchers from UCL and NYU (Lewis et al., 2020). We explain the three steps (retrieval, augmentation, generation) in how RAG makes AI smarter.

Classic RAG follows a fixed path: the question becomes a query, the search returns the top few passages and the model writes an answer. That works well when the answer sits in one place and the question is phrased much like the document.

It starts to break down when:

  • the question has several steps ("compare the terms in contracts A and B and check which complies with the new policy"),
  • the first search returns too little or passages stripped of context,
  • the answer is scattered across several documents or systems,
  • the question is vague and needs clarifying or breaking into parts first.

Classic RAG does not know it found too little. It answers from whatever it got, and that is where half-truths come from.

How does agentic RAG work?

Agentic RAG turns search from a one-off step into a tool the agent uses as many times as it needs. The agent, a language model running in a loop (more in what is agentic AI), goes through these stages:

  1. Analyze the question. The agent works out what is really being asked and breaks a complex question into sub-questions.
  2. Plan the search. It picks a source and a method: semantic search, full-text search, a database query, sometimes an external API.
  3. Retrieve. It calls the tool and gets passages back.
  4. Evaluate. It checks whether the passages answer the question and whether they are complete and consistent.
  5. Correct. If not, it rewrites the query, pulls in wider context (a whole section instead of a paragraph), searches another source or checks an exact phrase.
  6. Answer with citations. Once the context is sufficient, it writes an answer where every claim points to a source.
  7. Check. Before it goes out, the answer is verified for grounding in the sources.

Research on this has been going on for several years. Self-RAG trains a model to decide on its own when to retrieve and to critique the passages it finds (Asai et al., 2023). Corrective RAG adds an evaluation of retrieval quality and corrective actions when results are weak (Yan et al., 2024). Agentic RAG combines these ideas with general agent patterns.

RAG vs agentic RAG: the key differences

Classic RAGAgentic RAG
FlowFixed sequence: question, retrieve, answerLoop: plan, retrieve, evaluate, correct, answer
Number of searchesOneAs many as needed (with a limit)
Who decides how to searchThe pipeline codeThe model, while it works
SourcesUsually one vector storeSeveral tools and sources
Complex questionsWeakStrong
Latency and costLowerHigher
PredictabilityHighLower, needs a step trace
Best forFAQs, simple documentation questionsMulti-document analysis, multi-step questions, research

A simple analogy: classic RAG is a cook following a recipe to the letter. Agentic RAG is an experienced chef who tastes along the way, adjusts ingredients and changes technique when something is off.

What agentic RAG looks like in practice: our knowledge base

Our agentic knowledge base is a working example of this approach. The knowledge base built on the public drug registry holds over 2 million chunks from 22,786 documents. Here is the path from document to answer:

  • Documents. PDFs, Excel sheets and scans go into one pipeline. OCR reads scans in Polish and English, tables become text with their column headers, and a vision model describes charts. Each chunk gets a prefix with the document and section name before it becomes a vector.
  • Hybrid search. Two searches run in parallel: vector search by meaning and full-text search that handles Polish inflection by reducing words to their stems. A reranker orders results by relevance, and the two lists are merged with Reciprocal Rank Fusion (RRF).
  • Agent loop. Search is a tool for the model: it decides what to call and whether it has enough context. Asked about dosage, it finds the first results too thin, pulls in the full dosage section, then checks the interval between doses in the leaflet with an exact phrase.
  • Sourced answer. Every claim cites a specific passage in a document. Before it goes out, the answer passes guardrails: a grounding check against the sources and a hallucination detector.
  • Observability. Every agent step is recorded in the observability layer, so you can trace where the answer came from.

Everything runs on-premise on our own GPUs, so documents and questions never leave the company. That matters in finance, insurance and pharma, where we deploy these knowledge bases. We show what AI on your own infrastructure looks like on our AI infrastructure page.

Two lessons from this work. First, an agent will not fix poor data quality: if a scan is unreadable or a table has lost its headers, no number of iterations will recover it. Second, prefixing each chunk with the document and section name is a simple technique that noticeably helps retrieval. Anthropic described a similar approach (contextual retrieval): adding context to chunks combined with full-text search cut failed retrievals by 49%, and by 67% with reranking, in its tests (Anthropic, 2024).

What does agentic RAG give a business?

Agentic RAG answers questions that classic RAG answers poorly or not at all. In practice, that means:

  • Better answers to complex questions. The agent combines information from several documents and checks details instead of answering from the first hits.
  • Fewer confident mistakes. When context is missing, the agent keeps searching or admits it did not find an answer.
  • Verifiability. Citations and a step trace let someone check an answer in seconds, which builds user trust.
  • Easier to add sources. A new source is a new tool for the agent. Tools can be connected through the open MCP standard.

Use cases: customer support grounded in documentation, contract and policy analysis, searching technical documentation and procedures, and building reports from multiple sources.

When is classic RAG enough?

When questions are simple, the answer lives in one place and response time matters. FAQs, policy lookups, opening hours or the returns procedure are handled faster and cheaper by classic RAG.

With a very small knowledge base, you may not need RAG at all. Anthropic notes that a knowledge base under 200,000 tokens (about 500 pages) can simply be included in the prompt in full (Anthropic, 2024).

SituationRecommendation
A few dozen pages, simple questionsWhole knowledge base in the prompt, no RAG
Large knowledge base, single-fact questionsClassic RAG with hybrid search
Large knowledge base, complex or multi-step questionsAgentic RAG
Several systems and actions after the answerAgentic RAG with tools and human approval

What are the risks and how do you reduce them?

More autonomy means more places where things can go wrong. The main risks and how to handle them:

  • Latency and cost. Every iteration is another model call. Set a step limit and a per-question budget.
  • Endless loops. An agent can keep searching forever. You need a stopping condition, with "not found" accepted as a valid result.
  • Permissions. The agent must not see documents the person asking cannot access. Access control has to work at the search level, not just in the interface. We cover this in secure RAG data access.
  • Prompt injection. A document in the knowledge base can contain hidden instructions. The agent should treat document content as data, not commands.

How do you get started with agentic RAG?

Start with solid classic RAG and a test set, then add the agent loop where tests show that one search is not enough.

  1. Collect a few dozen real questions with correct answers and sources. That is your baseline.
  2. Get document processing right: OCR, tables, chunking with document and section context.
  3. Run hybrid search with reranking. For documents in inflected languages such as Polish, full-text search has to handle word forms.
  4. Add the agent loop with a step limit and measure how much it improves answers to hard questions.
  5. Enforce citations and a grounding check, and log every step.

Agentic RAG is the natural next step after RAG: the system stops just retrieving and starts checking whether what it found is enough. Important decisions still stay with people, but people get an answer they can verify in seconds.

Sources