Eval-Driven Development for Prompts
If you change a prompt without a test set, you are debugging in production. Here is how to build a prompt eval loop you will actually run.

The single highest-leverage habit in prompt work is not a clever technique. It is having a test set — a fixed list of inputs and the outputs you consider correct enough — so that every prompt change is a measurable bet instead of a vibe check.
Build the smallest set that catches regressions
You do not need thousands of examples. You need maybe 30-100 inputs that cover the failure modes you have actually seen. Add a case every time a bug reaches you and never remove one — the set only grows.
- Include the obvious happy path so you notice when you break it.
- Include the adversarial cases that previously produced hallucination.
- Include edge inputs: empty, huge, ambiguous, multi-lingual.
Grade automatically, review manually
Use an LLM as a judge for pass/fail on each case, but read the failures yourself. A judge that agrees with you 85% of the time is plenty — its job is to surface which cases moved, not to be the oracle.
for case in test_set:
out = run_prompt(case.input)
passed = judge(case.input, out, case.rubric)
if not passed:
print(case.id, "->", out[:120])
Treat changes like commits
Before you ship a prompt edit, run the set. If the score drops, do not ship — or explicitly decide to trade one failure for another and note it. This turns prompt changes from gambles into decisions.
A prompt without a test set is a hypothesis. A prompt with a test set is a product.



