Skip to main content

Agentic AI & RAG

How to stop chunks from losing meaning

Retrieval quality lives or dies at the chunk boundary. Overlap, metadata, and the right chunk size keep meaning intact.

CAPTIVOLT INSIGHTS · Published · Updated · 4 min read

Executive summary

Most RAG quality problems trace back to how documents were split. Chunk too small and the model loses the context to synthesise an answer; chunk too large and the real answer gets lost in the middle while the vector match blurs. Getting chunking right (overlap, metadata, and a size matched to your embedding model and content) is one of the highest-leverage things you can do for retrieval quality.

The problem

Splitting a document into naive fixed-size chunks breaks meaning at the boundaries: a clause is severed from the condition it depends on, a table is split across chunks and becomes unreadable, an answer is clipped halfway. The embedding model then indexes fragments that no longer mean what the author intended, and retrieval quietly degrades in ways demos never reveal.

A practical framework

  1. 01

    Use a sliding window. Configure a chunk overlap of roughly 10–20% (for example 50–100 tokens on a 512-token chunk) so context carries across boundaries.

  2. 02

    Enrich with metadata. Store the parent document, section title, and adjacent chunk IDs on each vector, so the system can pull surrounding context when an answer is clipped.

  3. 03

    Match chunk size to the embedding model. The chunk must fit the model’s input limit: a 256-token model silently ignores anything past 256, however large you set the chunk.

  4. 04

    Balance noise against context. Around 256–512 tokens is the general benchmark: too small and the model hallucinates for lack of context, too large and the answer is lost in the middle while the embedding turns generic.

  5. 05

    Adapt to the content. Tables → smaller chunks wrapped in structural markers; legal and financial → larger chunks (512–768) so conditional clauses stay intact; support knowledge bases → small chunks (128–256) for specific questions.

  6. 06

    When in doubt, use parent-child. Search on small child chunks for pinpoint accuracy, then return the larger parent chunk to the model for context: accurate retrieval and rich context together.

Experimental design

The method, published before the data.

Status: method published, results pending.

What this is, and is not

No results yet, and none are shown. This is the method, published before the data. The figures in the framework above (overlap of 10–20%, chunks of 256–512 tokens, 512–768 for legal text) are heuristics widely used in practice, not outputs of this experiment. The experiment is how they would be tested, rather than repeated.

Hypotheses

H1 · Overlap

An overlap of 10–20% answers boundary-spanning questions more accurately than no overlap, without reducing precision on pinpoint questions.

H2 · Size by content type

The best-performing chunk size differs by content type: larger for contracts, where a clause has to stay with its conditions, smaller for question-and-answer knowledge bases.

H3 · Parent-child

Searching small child chunks and returning their parent answers context-dependent questions more accurately than any single fixed size.

Method

Corpus

A mixed corpus containing the four content types the article names (contracts, tables, procedures and support articles), so a result for one type is never reported as a result for all.

Question set

Questions with agreed answers, of three kinds: boundary-spanning, pinpoint, and table lookups. Written before any configuration is run.

Variables

Chunk size, overlap and strategy (fixed, sliding window, parent-child), varied one at a time.

Held constant

The embedding model, the reranker, the generator and the number of passages retrieved, so that any difference is attributable to chunking.

Measures

Retrieval precision on the top passages, claim-level groundedness of the answer, and accuracy on boundary-spanning questions reported separately from the rest.

What would refute each

H1 is refuted if

Overlap does not improve boundary-spanning accuracy, or improves it only at the cost of pinpoint precision.

H2 is refuted if

A single chunk size performs best across all four content types.

H3 is refuted if

A single fixed size matches parent-child on context-dependent questions.

Running it, and publishing the data with the corpus description and the question set, is what would move this piece into Captivolt Research. Until then it stays a field note, which is what it is.

Practical implications

  • Most RAG failures are chunking failures.
  • Overlap and metadata help preserve meaning at the boundaries.
  • Chunk size must fit the embedding model’s token limit.
  • Content type dictates strategy: tables, contracts, and FAQs each want different sizes.
  • Parent-child chunking gives precise search with rich context.

About this article

Author
Captivolt Insights
Published
· updated

References

  1. Lost in the Middle: How Language Models Use Long Contexts · Liu et al., arXiv:2307.03172