What Is RAG? Retrieval-Augmented Generation Explained
RAG — Retrieval-Augmented Generation — fetches relevant information at question time and gives it to a language model to answer from, rather than relying on what the model learned during training.
It's the standard way to make a model answer from knowledge that's private, fresh, or citable — closing the gap between what a model knows and what you need it to know, without retraining.
The problem RAG solves
A language model is trained once, on a snapshot of data, then frozen. That leaves three hard limits: it knows nothing after its training cutoff, nothing private, and — worst — when asked something it doesn't know, it tends to invent a fluent, confident, wrong answer instead of admitting the gap. That last failure is called hallucination.
RAG fixes all three at once. Instead of hoping the answer is baked into the model's weights, you retrieve the relevant information at question time and hand it to the model with the question. The model answers from what you gave it.
The five-step pipeline
Every RAG system, however elaborate, is the same five steps. Learn them once and every variation reads as a modification of these five.
- Chunk — split documents into passages, small enough to be precise and large enough to be meaningful.
- Embed — convert each chunk into a vector capturing its meaning, and store it.
- Retrieve — embed the question the same way and find the closest chunks.
- Rerank — reorder the retrieved chunks so the best land on top, drop the rest.
- Generate — put the top chunks and the question into a prompt, and answer from them.
# indexing (once, ahead of time)
chunks = chunk(load(documents))
index.add(embed(chunks))
# querying (per question)
q = embed(question)
hits = index.search(q, top_k=20)
top = rerank(question, hits)[:5]
answer = llm(prompt(question, top))
Two phases: indexing and querying
RAG has two phases that are easy to mix up. Indexing happens offline, ahead of time: chunk your documents, embed them, store them. Querying happens live, per question: embed the question, retrieve, rerank, generate. They share the embedding model — the question must be embedded the same way the chunks were — but otherwise they're separate systems with separate concerns.
Why it beats hoping the model knows
The whole design goal is to make the model behave like a careful reader summarizing the passages in front of it — not an expert answering from memory. That shift is what makes answers trustworthy: you can update the knowledge, control who sees what, and cite the source of every claim. None of that is possible when the knowledge is locked inside the model's weights.
RAG turns "answer from memory"
into "answer from these documents."
Where RAG fits
RAG is one layer of modern AI engineering. It's the retrieval slice of context engineering — deciding what enters a model's context window. It becomes a tool that agents call. And when retrieval is exposed as a tool, protocols like MCP and A2A are how agents reach it and each other. Master the pipeline and you hold the piece the whole stack depends on.