All articlesPrompt Engineering

Eval-Driven Development for Prompts

If you change a prompt without a test set, you are debugging in production. Here is how to build a prompt eval loop you will actually run.

Sri Raman18 August 20269 min read
Eval-Driven Development for Prompts

The single highest-leverage habit in prompt work is not a clever technique. It is having a test set — a fixed list of inputs and the outputs you consider correct enough — so that every prompt change is a measurable bet instead of a vibe check.

Build the smallest set that catches regressions

You do not need thousands of examples. You need maybe 30-100 inputs that cover the failure modes you have actually seen. Add a case every time a bug reaches you and never remove one — the set only grows.

  • Include the obvious happy path so you notice when you break it.
  • Include the adversarial cases that previously produced hallucination.
  • Include edge inputs: empty, huge, ambiguous, multi-lingual.

Grade automatically, review manually

Use an LLM as a judge for pass/fail on each case, but read the failures yourself. A judge that agrees with you 85% of the time is plenty — its job is to surface which cases moved, not to be the oracle.

for case in test_set:
    out = run_prompt(case.input)
    passed = judge(case.input, out, case.rubric)
    if not passed:
        print(case.id, "->", out[:120])

Treat changes like commits

Before you ship a prompt edit, run the set. If the score drops, do not ship — or explicitly decide to trade one failure for another and note it. This turns prompt changes from gambles into decisions.

A prompt without a test set is a hypothesis. A prompt with a test set is a product.
Share this article