Prompt LibraryAgents

Agent Benchmark Task Suite Designer Prompt

Design a benchmark suite to compare agent frameworks.

The prompt

Design a benchmark suite to compare agent frameworks (e.g., CrewAI, AutoGen, LangGraph).
10 tasks across: single-tool, multi-tool, tool-selection, error-recovery, and human-escalation.
For each task: the setup, the success criterion, the budget (max steps, max cost), and what a "cheating" solution would look like (so it can be rejected).
Ensure tasks are model-agnostic (don't require a specific model's quirks).

Replace the {{fields}} with your own context and tighten the rules to match your domain.

Share this prompt