All articlesLLM Ops

Retrieval Reranking That Actually Works

Vector search gets you candidates. Reranking decides which ones the model sees. Here is how to build a rerank step that earns its latency.

Sri Raman22 August 20267 min read
Retrieval Reranking That Actually Works

Vector retrieval is cheap and good enough to fetch candidates. It is also bad at fine distinctions — the top result and the tenth result are often nearly tied in cosine distance, but only one of them answers the user's actual question. Reranking is the step that fixes that, and most teams either skip it or bolt on a model without measuring it.

Why raw similarity fails

Embedding similarity measures topical closeness, not answerhood. Two passages can be about the same entity while one describes the thing you asked and the other contradicts it. A reranker, which reads the query and the passage together, can tell the difference.

A rerank pipeline worth keeping

  1. 1Retrieve top-k=50 with the vector index — cheap, recall-oriented.
  2. 2Score each candidate with a cross-encoder reranker on (query, passage).
  3. 3Keep top-n=5 after reranking, deduplicate by semantic near-duplicates.
  4. 4Only then compose the context window.

The latency cost is one extra model pass over 50 short pairs — usually acceptable. The accuracy lift is almost always larger than any tweak to the prompt.

Measure before you ship

Hold out a set of queries where you know the right passage. Report recall@5 before and after reranking. If reranking does not move recall@5 by a meaningful amount, the model is not earning its latency — drop it or swap it.

# the only metric that matters early
recall_at_5 = len(set(pred_top5) & set(correct)) / len(correct)
Retrieval is a recall problem. Reranking is a precision problem. Conflating them is why most RAG feels mushy.
Share this article