Agent Observability: Logging, Tracing, and Alerting in Practice
An agent you cannot observe is an agent you cannot debug. What to log, how to trace, and which alerts wake you up at 3am.

Observability for agents is not log messages — it is the ability to reconstruct what happened during a run and answer 'why did it fail?' An agent that calls five tools, replans twice, and escalates once generates a trace. Without that trace, debugging is archaeology.
The trace is the unit of observability
Every agent run gets a trace ID. Every tool call, every replan, every escalation is logged under that ID with a timestamp. The trace is the story of the run. When a user reports 'the agent did the wrong thing,' you pull the trace and see exactly which step went wrong — not a guess.
What to log at each step
For each tool call: the input (hashed, not raw, if it contains PII), the output (hashed), the latency, and the status. For each replan: the trigger reason and the revised plan. For each escalation: the confidence score and the guardrail that fired. Do not log raw user text — hash it. Do not log full tool outputs — hash them. You need to correlate runs, not reconstruct conversations.
The metrics that matter
Success rate, p50/p95 latency, cost per run, tool-call error rate, loops-to-completion, and guardrail-trigger rate. These six metrics tell you whether the agent is healthy. Success rate alone hides cost spirals; cost alone hides quality drops. You need all six on one dashboard.
The one alert that matters
Alert on guardrail-trigger rate, not success rate. A spike in guardrail violations means the agent is doing something unsafe, and that is the 3am problem. A dip in success rate is a quality problem — investigate it in business hours. Do not wake someone up because the agent got 4% less accurate; wake them up if it tried to do something forbidden.



