How to Estimate RAG Costs Before You Build

How to Estimate RAG Costs Before You Build

How to Estimate RAG Costs Before You Build

RAG has two cost centers: a one-time indexing cost (embedding all your chunks) and a recurring per-query cost (embedding the question plus sending retrieved chunks into generation).

Both are driven by tokens, so once you know your chunk token counts and your model's rates, you can estimate the bill before you build — and see which levers actually move it.

The two cost centers

RAG cost splits cleanly in two. Indexing is a one-time (or occasional-update) cost: you embed every chunk in your corpus once so it can be retrieved later. Querying is the recurring cost: for each question you embed the query, retrieve the top chunks, and send them into the generation model. They scale with different things, so it's worth estimating them separately.

Estimating indexing cost

Indexing cost is the total tokens across all your chunks times your embedding model's rate. It's usually the smaller of the two for a system that answers many queries, because you pay it once. But for a large corpus it's real, and overlap inflates it — every overlapping token is embedded more than once.

index_cost = total_chunk_tokens / 1_000_000 * embedding_rate
# overlap inflates total_chunk_tokens: repeated text is embedded twice

Estimating per-query cost

Per-query cost is dominated by generation: you send the retrieved chunks (top-k of them) into the model along with the question. More chunks or larger chunks means more tokens per query, at the generation rate — which is usually far higher than the embedding rate. This is the cost that recurs on every single query, so it's the one that adds up.

Indexing you pay once.
Generation you pay every query.
Want to see this on your own text? The free RAG Chunk Visualizer shows your chunks, token counts, and quality flags right in the browser.Try the Free Chunk Visualizer

The levers that actually move the bill

  • Top-k — how many chunks you send into generation. The most direct lever on per-query cost.
  • Chunk size — larger chunks mean more tokens per retrieved item.
  • Overlap — inflates indexing cost by embedding repeated text.
  • Model tier — generation rates vary widely; the same pipeline can cost 5–15x more on a premium model.

See the cost before you commit

You don't have to build the pipeline to know what it'll cost. The RAG Chunk Visualizer estimates both indexing and per-query cost from your actual chunks, across budget, standard, and premium model tiers, and lets you turn the top-k and overlap knobs to watch the bill respond. Cost-model presets and top-k modelling are part of the Full Edition.

The Full Edition unlocks overlap control, cost-model presets, top-k modelling, strategy comparison, JSON export, and vector-DB record preview — an installable web app.Get the Full Edition