Prompt LibraryEngineering

LLM Evaluation Dataset Builder Prompt

Create a small honest eval set for a prompt or agent.

The prompt

Task the system performs: {{task}}
Write 12 evaluation cases: 6 typical, 4 edge, 2 adversarial.
For each: the input, the pass criteria in one checkable sentence, and why it is included.
Prefer cases that would fail today over cases that flatter the system.

Replace the {{fields}} with your own context and tighten the rules to match your domain.

Share this prompt