All articlesPrompt Engineering

Few-Shot Engineering: Curating Examples That Actually Teach

The right three examples beat thirty mediocre ones. A practical guide to selecting, ordering, and stress-testing few-shot demonstrations for reliability.

Sri Raman29 August 20269 min read
Few-Shot Engineering: Curating Examples That Actually Teach

Few-shot prompting is the most overused and least understood technique in prompt engineering. Teams paste a dozen examples into a prompt, see the output improve, and conclude that more examples are always better. They are not. More examples mean more tokens, slower responses, and often, more confusion when examples contradict each other.

Diversity beats quantity

The goal of few-shot examples is to teach the model the input-to-output mapping. If all your examples look the same, the model learns one pattern and applies it everywhere — including where it does not fit. Pick examples that cover different input shapes: a typical case, an edge case, and a boundary case. Three diverse examples will generalize better than ten near-duplicates.

Order matters more than you think

Models are influenced by recency. The last example before the real input carries the most weight. Put your most representative example last, and your trickiest edge case earlier. This is not folklore — it shows up in ablation studies where reordering the same examples shifts accuracy by several points.

The anti-pattern: examples that teach the wrong rule

If every example output starts with 'Sure!' the model learns that the correct answer always begins with 'Sure!'. If every example includes an apology, the model apologizes in production. Audit your examples for accidental patterns the model will over-generalize.

When to stop adding examples

Run your prompt with zero, one, three, and five examples on the same test set. Plot accuracy. The curve usually flattens between three and five. Beyond that, you are paying for tokens that do not improve output. If five examples do not solve the problem, the issue is the task decomposition or the instruction, not the example count.

Few-shot examples are a teaching tool, not a brute-force lever. If the model needs twenty examples to get it right, your instruction is under-specified.
Share this article