Evaluating Autonomous Agents: Beyond Single-Turn Metrics
Single-turn evals do not capture agent behavior. How to measure success rate, cost-per-success, and safety for multi-step agent loops.

Evaluating an LLM response is hard enough. Evaluating an agent that takes ten steps, calls four tools, and may replan twice is an order of magnitude harder. The metrics that work for chat — BLEU, human preference, accuracy on a benchmark — do not transfer.
The metrics that matter for agents
A useful agent eval suite measures five things: success rate (did it achieve the goal?), cost per successful run (how much did the wins cost?), step efficiency (how many steps for a task that should take three?), safety violation rate (how often did it cross a guardrail?), and unnecessary-tool-call rate (how often did it call a tool it did not need?).
Cost per success is the north star
Two agents can both achieve 90% success rate. One costs $0.02 per run, the other $0.50. If you optimize only success rate, you pick the expensive one. Cost per success — total spend divided by successful runs — captures the trade-off. It also penalizes agents that retry endlessly: a 95% success rate that burns $5 per win is worse than a 90% rate at $0.10.
Designing tasks that test the right thing
Your eval tasks must include: single-tool tasks (can it call the right tool?), multi-tool tasks (can it sequence?), error-recovery tasks (can it replan when a tool fails?), and escalation tasks (does it know when to ask a human?). If all your tasks are single-tool, you are testing function calling, not agency.
Detecting infinite loops
An agent that calls the same tool with the same arguments three times in a row is stuck. Your harness must detect this and kill the run, recording it as a failure. Without this, a bug where the agent loops forever will inflate cost without showing up as a success-rate problem.


