All articlesLLM Ops

RAG Pipeline Architecture: From Naive to Production

A naive RAG pipeline retrieves and stuffs. A production pipeline chunks, reranks, grounds, and guards. Here is the difference.

Sri Raman25 August 202611 min read
RAG Pipeline Architecture: From Naive to Production

The simplest RAG pipeline: embed the query, retrieve top-k chunks, stuff them into the prompt, generate. It works for a demo. It fails in production because retrieval is noisy, chunks are the wrong size, and the model hallucinates over gaps.

Chunking is not a rounding error

Fixed-size chunking with an overlap is the default, but it breaks semantic boundaries — a chunk can split a paragraph mid-sentence. Semantic chunking (by heading or paragraph) preserves meaning but produces variable-size chunks that complicate embedding. The pragmatic middle ground: chunk by heading with a max-size cap, and store the heading path as metadata for filtering.

Reranking is where the quality lives

Vector retrieval is fast but imprecise — it returns chunks that are semantically similar, not necessarily relevant. A cross-encoder reranker, applied to the top 20-50 retrieved chunks, re-scores them for true relevance and cuts to top 5-10. This single step typically improves answer quality more than any other change.

Grounding checks after generation

After the model generates, verify every claim is supported by the retrieved chunks. If a claim has no supporting chunk, either remove it or mark it as '[not in sources]'. This turns RAG from a retrieval-stuffing hack into a grounded system. Users trust answers they can trace.

Handle zero-result queries

When retrieval returns nothing relevant, do not let the model answer from its parametric memory. Return 'I don't have information on that.' The worst RAG failure is a confident hallucination over an empty retrieval.

Share this article