Reranking in RAG: Retrieve Wide, Rerank Narrow
A reranker is a second, more accurate model that reorders retrieved chunks by reading the query and each chunk together and scoring how well the chunk answers the query.
The pattern is retrieve wide, rerank narrow: fetch many candidates with fast vector search, then use the reranker to float the best few to the top and send only those to the model. It's one of the highest-return additions to a RAG pipeline.
Why retrieval alone isn't enough
The embedding search that retrieves your candidates is optimized for speed across millions of chunks, and that speed costs precision: the top result by vector distance is often not actually the most relevant. A reranker fixes this. It looks at the query and each candidate chunk together and scores how well the chunk answers the query, then reorders them.
Why a reranker is more accurate
The difference is in how they compare. The embedding search compared two vectors computed separately — the query never "saw" the chunk. A reranker, a cross-encoder, reads the query and the chunk at the same time, so it judges relevance directly rather than by proxy. That joint reading is far more accurate, and far too slow to run over your whole corpus — which is exactly why you run cheap retrieval first and expensive reranking only on the survivors.
The retrieve-wide, rerank-narrow pattern
# 1. retrieve generously with fast vector search
candidates = index.search(embed(question), top_k=25)
# 2. rerank precisely with a cross-encoder
scored = reranker.score(question, candidates)
# 3. keep only the best few for the prompt
top = sort_by_score(scored)[:5]
Retrieval finds candidates fast and blunt.
Reranking orders them slow and sharp.
It also cuts your cost
Reranking does more than improve ordering — it lets you send fewer, better chunks to the model. Rather than stuffing twenty mediocre chunks into the prompt (slow, expensive, and prone to "lost in the middle"), you send the five the reranker is confident about. And generation tokens are usually the largest cost in the pipeline, so cutting the prompt from twenty chunks to five cuts your biggest cost on every query — while improving the answer. The reranker pays for itself.
When you can skip it
Reranking is high-return but not free, and a few situations don't need it. If your corpus is small and questions map cleanly to single chunks, plain retrieval may already put the right passage first. If latency is critical and quality is already acceptable, the extra stage may not earn its milliseconds. Add the reranker when evaluation shows retrieval ordering is your bottleneck, and skip it when the numbers say you don't need it.