Building a production RAG pipeline: the components that matter
A RAG demo can be built in an afternoon with a document loader, an embedding call, and a vector store. A production RAG pipeline requires several additional components that determine whether the system remains accurate and maintainable as documents, users, and query volume grow. This post walks through those components in order.
Ingestion and parsing
Document parsing quality sets a ceiling on everything downstream. PDFs with tables, scanned documents requiring OCR, and structured formats like spreadsheets each need parsing logic tailored to preserve their structure, rather than a single generic text extractor applied to all formats.
Build ingestion as an idempotent pipeline that can safely reprocess a document that changes, rather than a one-time script. Track a content hash per document so unchanged documents are skipped and changed documents trigger re-chunking and re-embedding automatically.
Chunking, embedding, and indexing
Chunking strategy should be tuned to the document types identified during ingestion rather than applied uniformly across all content. Store chunk metadata, including source and section, to support both filtering and citation at generation time.
Choose a vector index that supports metadata filtering natively, since most production queries benefit from narrowing the search space by document type, date, or access permission before similarity search runs, rather than filtering after retrieval.
Retrieval and reranking
Pure vector similarity search often retrieves chunks that are topically related but not the most relevant to the specific query. Adding a reranking step, using a smaller cross-encoder model to reorder the top candidates from initial retrieval, consistently improves the relevance of what reaches the generation step.
Hybrid retrieval, combining vector similarity with keyword or lexical search, catches cases where a query includes specific terms, identifiers, or names that embedding similarity alone handles poorly. This combination is worth the added complexity for most enterprise document sets.
- Retrieve a wider candidate set with vector search, then narrow with reranking
- Combine lexical and vector search for queries containing specific identifiers or names
- Enforce document-level access control before results reach the generation step
Generation with grounded prompting
The prompt sent to the generation model should clearly separate retrieved context from the instruction, and should instruct the model explicitly on how to behave when the retrieved context does not contain a relevant answer, rather than allowing it to fall back on general knowledge silently.
Include citations back to source chunks in the generated output where the application supports it. This both improves user trust and makes it possible to trace an incorrect answer back to the retrieval step or the source document that caused it.
Monitoring and continuous improvement
Log queries, retrieved chunks, and generated answers together, so quality issues can be traced to the specific stage responsible. Periodically sample this log for review against a quality rubric, and use recurring failure patterns to guide chunking, retrieval, or prompt adjustments.
Treat the pipeline as something that improves continuously based on production feedback, not something finalized at launch. The most durable production RAG systems we have seen invest ongoing effort into refining retrieval based on real query patterns rather than the patterns anticipated during initial design.
Want this applied to your data?
We scope private AI projects in one call and start with a pilot you can evaluate.