Agentic RAG: When the Model Decides What to Retrieve

Agentic RAG: When the Model Decides What to Retrieve

Agentic RAG: When the Model Decides What to Retrieve

Agentic RAG turns retrieval into a tool the model calls on demand — deciding whether to retrieve, what query to search with, and whether the results are enough — instead of running retrieval once as a fixed step.

It handles open-ended, multi-step questions that a single retrieval can't, at the cost of more latency and less predictability. It's where RAG meets the wider world of AI agents.

From a fixed step to a tool

In classic RAG, retrieval happens exactly once, before generation, whether or not it was needed. Agentic RAG hands the decision to the model itself: retrieval becomes a tool the model can choose to call, with a query it writes, as many times as the question demands. The mental shift is small and total — retrieval stops being step three of a pipeline and becomes something the model invokes when it decides it needs to.

What the model gets to decide

  • Whether to retrieve at all — a greeting or simple reasoning question needs no lookup; the model can skip it.
  • What to search for — the model writes the query, often several, rather than searching with the raw question verbatim.
  • Whether it has enough — after reading results, the model can judge them insufficient and retrieve again with a refined query.

The self-correction loop

The most useful agentic pattern is the simplest: after retrieving, the model checks whether the results actually answer the question, and if not, reformulates and retrieves again. This one loop catches a large share of the failures that sink classic RAG — the near-miss retrieval, the ambiguous query, the question that needed rephrasing. It costs an extra round-trip when it triggers and buys a meaningful lift on hard questions. If you add one agentic behavior, add this one.

Classic RAG retrieves, then answers.
Agentic RAG retrieves until it can answer.
Want the whole pipeline on a few pages — chunk, embed, retrieve, rerank, generate, and the decisions that matter? Grab the free RAG Quick-Start.Download Free — RAG Quick-Start

The cost of the flexibility

Agentic RAG is more capable and less predictable. Each extra retrieval is latency and money, the loop can wander, and a model that decides not to retrieve when it should is a new failure mode. The trade is real: classic RAG is cheaper, faster, and easier to reason about; agentic RAG handles open-ended, multi-step questions a single retrieval can't. Start classic, and move to agentic when the questions genuinely need it — not because it sounds more advanced.

Where it connects to the stack

Agentic RAG is the bridge between retrieval and the wider world of AI agents. Retrieval-as-a-tool is exactly the kind of capability an agent reaches through a protocol like MCP; a retrieval agent is exactly the kind of specialist another agent might delegate to over an agent-to-agent protocol like A2A. RAG stops being a standalone pipeline and becomes one capability in an agent's toolkit.

A2A: The Complete Guide to the Agent2Agent Protocol is the full reference — 46 pages, 15 chapters, 7 appendices, with a worked example and a 30-day adoption path.Get the Complete Guide