RAG Chunking Strategies: The Decision That Caps Your Quality
Chunking splits your documents into the passages RAG retrieves and feeds to the model. Its size and boundaries shape everything downstream.
The highest-return rule: chunk on the document's structure — headings, sections, paragraphs — not blindly by a fixed number of characters. Add modest overlap, and keep each chunk's metadata.
Why chunking decides your ceiling
Before anything can be retrieved, documents have to be broken into pieces — chunks — because you retrieve and feed passages, not whole documents. The size and boundaries of those chunks shape everything after. Retrieval, reranking, and generation can only work with the chunks you gave them, so a chunking mistake caps the quality of the entire system no matter how good the later stages are.
The core trade-off
Chunk size is a tension between precision and context. Small chunks retrieve precisely — you get exactly the sentence that matters — but that sentence may be meaningless without its surroundings. Large chunks preserve context but retrieve bluntly, dragging in unrelated material that dilutes the answer. There's no universal right answer; it depends on your documents and your questions.
A common starting point is a few hundred tokens per chunk, then tuning from there against evaluation.
Overlap, and why it helps
Chunk overlap means each chunk repeats the last stretch of the previous one. It costs a little storage and prevents a specific failure: an answer that straddles a chunk boundary, half in one piece and half in the next, so neither chunk alone contains it. A modest overlap — ten to twenty percent — keeps boundary-spanning facts intact.
Respect the document's structure
The best chunk boundaries are the ones the document already has. Splitting on headings, paragraphs, or sections keeps semantically whole units together, which retrieves far better than splitting blindly every N characters. A chunk that is one complete subsection beats a chunk that ends mid-sentence every time.
# naive: split every N characters (blunt)
chunks = [text[i:i+800] for i in range(0, len(text), 700)]
# better: split on structure, then size-bound
chunks = split_on_headings_then_pack(text, target=800, overlap=100)
The document is the raw material.
The questions are the specification.
Chunk for the question, not the document
The subtle art is that you're really chunking for the questions you expect, not for the shape of the document. A reference manual answered with pinpoint factual questions wants small, precise chunks. A narrative answered with questions about reasoning wants larger chunks that preserve the argument. When you can't decide, look at real user questions — they tell you how much context an answer needs.
Common chunking mistakes
- Fixed-size everything — splitting blindly by character count, ignoring structure. The most common mistake and the easiest fix.
- No overlap — facts that straddle a boundary end up in neither chunk.
- Chunks too large — one chunk covers several topics, retrieves for one, drags the others along.
- Losing the metadata — chunking away the source, section, and date, then being unable to cite or filter later.
- Never re-chunking — treating the first attempt as final instead of a parameter to tune against evaluation.
Test your chunking before you trust it
Chunking is easy to get subtly wrong and easy to check. Take a handful of real questions, retrieve against your chunked corpus, and read the chunks that come back. Are they whole thoughts, or do they start mid-sentence? Does each answer sit cleanly inside a chunk, or is it split across two? Five minutes reading actual retrieved chunks tells you more than any amount of theory.