Chunking and embedding strategies that actually improve RAG
Retrieval quality in a RAG system depends more on chunking and embedding decisions than on the choice of generation model, yet these decisions are often made once early in a project and never revisited. This post covers the tradeoffs we see teams get wrong most often when building retrieval pipelines.
Chunk size is a tradeoff, not a default
Small chunks improve retrieval precision because each chunk is more topically focused, but they can strip away context a generation model needs to answer correctly. Large chunks preserve context but dilute the embedding signal and can push relevant content below the retrieval threshold when mixed with irrelevant surrounding text.
There is no universal correct chunk size. Technical documentation with dense, self-contained paragraphs often works well with smaller chunks, while narrative or legal text with cross-referencing content often needs larger chunks or explicit context injection.
Structure-aware chunking beats fixed-length splitting
Splitting text at a fixed token count without regard to document structure frequently cuts a chunk in the middle of a table row, a numbered list, or a heading, producing embeddings that represent an incoherent fragment. Chunking along natural boundaries, such as headings, paragraphs, or table rows, generally produces better retrieval results for the same average chunk size.
For structured formats like PDFs with tables or code with function boundaries, invest in format-specific parsing before falling back to generic text splitting. The upfront effort pays off directly in retrieval quality.
Overlap and metadata
A modest overlap between adjacent chunks, often in the range of ten to twenty percent of chunk length, reduces the chance that a relevant sentence gets split across a boundary and misses retrieval entirely. Excessive overlap increases index size and storage cost without proportional quality gains.
Attach metadata to each chunk, including source document, section heading, and last-updated date. This metadata supports filtering at query time and lets the generation step cite sources accurately, which matters for user trust in the answer.
- Store source, section, and date metadata alongside every chunk
- Use overlap in the ten to twenty percent range as a starting point, then tune against your evaluation set
- Re-chunk when document structure changes significantly, not only when content changes
Choosing an embedding model
Embedding model choice affects both retrieval quality and infrastructure cost, since embeddings must be computed for every chunk in the corpus and for every query at inference time. Evaluate candidate embedding models against your own document set and query patterns rather than relying solely on public benchmark rankings.
Multilingual corpora require an embedding model trained for multilingual retrieval; using an English-only model on mixed-language content produces degraded results that are easy to miss until a non-English query fails silently.
Re-evaluate as the corpus grows
Chunking and embedding decisions that work well for a corpus of a few thousand documents can behave differently once the corpus grows by an order of magnitude, as retrieval competition increases and near-duplicate content becomes more common. Periodically re-run retrieval evaluation against a fixed query set as the corpus scales.
Build the pipeline so that re-chunking and re-embedding the corpus is a repeatable, low-effort operation. Teams that treat the initial chunking strategy as permanent tend to accumulate retrieval quality problems that are expensive to diagnose later.
Want this applied to your data?
We scope private AI projects in one call and start with a pilot you can evaluate.