Prompt LibraryEngineering
LLM Evaluation Dataset Builder Prompt
Create a small honest eval set for a prompt or agent.
The prompt
Task the system performs: {{task}}
Write 12 evaluation cases: 6 typical, 4 edge, 2 adversarial.
For each: the input, the pass criteria in one checkable sentence, and why it is included.
Prefer cases that would fail today over cases that flatter the system.Replace the {{fields}} with your own context and tighten the rules to match your domain.



