Observability for Agents: Logging the Right Thing
Agents fail in ways single calls cannot. To debug them you need traces of decisions, not just inputs and outputs. What to log.

A chat bot logs a prompt and a response. An agent logs a run that branches, calls tools, reads results, and revises. If your logging stops at the first and last call, you cannot debug the failure in the middle — which is where almost all agent failures live.
Log the trace, not the transcript
A transcript is what was said. A trace is what was decided: which tool was chosen and why, what it returned, whether the agent treated that as success or retry. The trace is the artifact you replay when something breaks.
- Every tool call: name, arguments, return, duration, and the model's stated reason for choosing it.
- Every loop iteration: what changed, what stayed, why it did not stop.
- The final answer plus the full decision path that produced it.
Cap and structure the trace
Raw traces balloon. Summarise tool outputs before logging them, keep a hard cap on iterations logged, and store traces as structured JSON so you can query 'runs that called search_docs more than three times'. A trace you cannot query is a trace you will never read.
{
"run_id": "...",
"steps": [
{ "type": "thought", "text": "Need the user's plan first." },
{ "type": "tool", "name": "get_plan", "reason": "free vs paid", "ms": 120 },
{ "type": "tool_result", "ok": true, "summary": "free tier" }
]
}
Replay is the real test
When a run fails, replay its trace against the new prompt or tool. If the failure disappears and nothing else breaks, you have a real fix. If you cannot replay, you are guessing — and with agents, guessing is expensive.
If you cannot replay a failed run, you do not have observability. You have a story about what probably happened.



