All articlesLLM Ops

Hallucination Reduction: Techniques That Actually Work

You cannot eliminate hallucination, but you can reduce it dramatically. Grounding, verification, and canaries — the techniques that survive production.

Sri Raman15 August 20269 min read
Hallucination Reduction: Techniques That Actually Work

Hallucination is not a bug you can fix with a prompt tweak. It is a property of how LLMs generate text — they predict the next token, and the most likely next token is not always the true one. You cannot eliminate it. You can reduce it, detect it, and design around it.

Grounding is the strongest defense

If the model answers from retrieved context, not parametric memory, hallucination drops sharply. The constraint: every claim must be traceable to a source. When the model cannot find a source, it must say 'I don't know.' This is RAG with a grounding check — and it is the single most effective hallucination reducer in production.

Chain-of-verification

After the model generates an answer, run a second pass that lists every factual claim and verifies each one. Claims that cannot be verified are struck. This doubles the token cost, but for high-stakes outputs (medical, legal, financial), the cost is justified. For low-stakes outputs (drafts, summaries), it is overkill.

The canary test

Plant a fictional entity in the system note — a detail the model should not know. If the output treats the fiction as fact, you have a hallucination. This is a canary: it does not prevent hallucination, but it detects it in a way that is unambiguous. A failed canary means the output needs review, not publication.

Temperature is not the lever people think

Lowering temperature reduces randomness, which slightly reduces hallucination. But a model that hallucinates at temperature 0.7 will hallucinate at temperature 0. The cause is not randomness; it is the model's uncertainty about the fact. Temperature is a minor knob; grounding and verification are the major ones.

Share this article