TL;DR: An agent failure is rarely a failed request. It is a bad decision made three steps earlier that only surfaced as a wrong answer. Traces that debug these failures need a span hierarchy that mirrors the agent loop (run, step, model call, tool call), attributes that explain decisions rather than just timing them, a deliberate policy for capturing prompt content, and tail-based sampling that keeps the interesting runs. Standard request tracing gives you none of this by default.
Why classic tracing falls short for agents
A conventional service trace answers "which hop was slow or errored?" That works because the call graph is fixed by code. An agent's call graph is chosen at runtime by a model. Two runs with the same input can take different paths, call different tools, and finish in different step counts. Anthropic's engineering team put it plainly when describing their multi-agent research system: agents "make dynamic decisions and are non-deterministic between runs, even with identical prompts," which makes debugging harder, and full production tracing is what let them diagnose failures systematically (How we built our multi-agent research system).
Three properties make agent failures different from ordinary service failures:
- Silent failures. The run returns HTTP 200 with a fluent, wrong answer. No span has
status=ERROR. Your error-rate dashboard stays green. - Causal distance. The root cause (a tool returned an empty list, the model treated it as "no results exist") sits several steps before the visible symptom.
- Compounding. A small mistake in step 2 changes the context for every later step. The same source article notes that minor system failures can be catastrophic for agents because errors compound.
So the design question is not "how do we log LLM calls?" It is "what would we need to see to reconstruct why the agent did what it did?"
The span hierarchy: mirror the loop, not the framework
Most agent runtimes are a loop: build context, call the model, parse a decision, maybe execute tools, repeat until a stop condition. Your trace should have the same shape.
invoke_agent (the whole run; one per user request)
├── plan / step 1
│ ├── chat (model call: decision + token usage)
│ ├── execute_tool (search_orders)
│ └── execute_tool (get_customer)
├── step 2
│ ├── retrieval (vector lookup, if you do RAG)
│ └── chat
├── step 3
│ └── chat (final answer; finish_reason=stop)
└── (guardrail / validation spans as siblings)The OpenTelemetry GenAI semantic conventions define roughly this vocabulary. Operation names include invoke_agent, invoke_workflow, plan, chat, execute_tool, retrieval and memory operations (Dash0's walkthrough of the conventions). The OpenTelemetry project's own GenAI write-up uses the same three-level shape: an invoke_agent root, chat children for each LLM call, and execute_tool spans for tool invocations (Inside the LLM Call: GenAI Observability with OpenTelemetry).
A "step" span is worth adding
The conventions do not require an explicit per-iteration span, but we recommend one as a custom grouping span. Without it, a run with 14 model calls and 22 tool calls is a flat list of 36 children. With it, "the agent looped 9 times" is visible at a glance, and you can attach a step index and the stop/continue decision to one place.
Pseudo-code for the loop
with tracer.start_as_current_span("invoke_agent", attributes={
"gen_ai.operation.name": "invoke_agent",
"gen_ai.agent.name": "support-triage",
"gen_ai.agent.version": AGENT_VERSION, # prompt+tool-set hash, see below
"gen_ai.conversation.id": conversation_id,
}) as run:
for step in range(MAX_STEPS):
with tracer.start_as_current_span("agent.step", attributes={"agent.step.index": step}) as s:
resp = call_model(messages) # emits a `chat` span
s.set_attribute("agent.step.decision", classify(resp)) # tool_call | final | clarify
if resp.tool_calls:
for tc in resp.tool_calls:
run_tool(tc) # emits `execute_tool` span
continue
break
else:
run.set_attribute("agent.stop_reason", "max_steps")
run.set_status(Status(StatusCode.ERROR, "step budget exhausted"))Note the last two lines. A run that hits its step budget is a failure even though no exception was thrown. Convert agent-level failure conditions into span status, otherwise tail sampling and alerting cannot see them.
Attributes that explain decisions
Timing and token counts are table stakes. The attributes that let you debug are the ones that record what the agent knew and chose.
| Question during an incident | Attribute / signal you need |
|---|---|
| Which version of the agent was this? | gen_ai.agent.version, plus a hash of system prompt + tool schemas + model ID |
| Did the model hit a limit? | gen_ai.response.finish_reasons (length vs stop vs tool_calls) |
| What did it cost? | gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, cache-read tokens |
| Which tool, with what call identity? | gen_ai.tool.name, gen_ai.tool.call.id |
| Why did the loop end? | custom agent.stop_reason: final_answer, max_steps, budget, guardrail, tool_error |
| Was the context truncated or summarized? | custom agent.context.tokens, agent.context.truncated |
| Was a tool result empty or malformed? | custom tool.result.size, tool.result.empty, tool.schema_valid |
| Which conversation does this belong to? | gen_ai.conversation.id |
The gen_ai.* names above come from the OpenTelemetry GenAI conventions; the agent.* and tool.* names are our own, and you should namespace your own attributes so they cannot collide with future standard ones.
Stability caveat: pin and centralize attribute names
The GenAI attributes are still at "Development" stability, and the conventions have moved to a dedicated repository for faster iteration (per the same Dash0 walkthrough). Different instrumentation libraries emit different dialects. Two practical mitigations: define attribute-name constants in one module of your own code, and normalize dialects at the collector (for example with the transform processor) rather than in dashboards.
Versioning is the highest-leverage attribute
When quality regresses on Tuesday, the first question is "what changed?" An agent's behavior is a function of model, system prompt, tool descriptions, retrieval index and temperature. Hash the prompt text and the serialized tool schemas into a single agent.config_hash, and stamp it on the root span. Then "failures started when hash a41f… rolled out" is a query, not an investigation. This also gives your LLM eval pipeline a join key between offline eval runs and production traces. If you are building that side, see our write-up on how to build an LLM eval pipeline from golden datasets to CI gates.
Content capture: the privacy and cost decision
Prompts, tool arguments and completions are what you actually read when debugging, and they are also where personal data, secrets and customer content live. The OpenTelemetry conventions treat this deliberately: prompt and completion content is excluded from default span attributes and exposed through opt-in attributes/events instead. As the OpenTelemetry post puts it for Copilot's instrumentation, by default no prompt content or tool arguments are captured because they can contain sensitive data (same OpenTelemetry post). Opt-in mechanics differ by library, so check yours.
A workable policy has three tiers:
- Always capture structure, never content: step counts, tool names, argument shapes (keys and types), result sizes, token counts, hashes, error types. This is cheap and safe, and it answers most "what path did the agent take?" questions. Anthropic describes taking this approach to production monitoring: tracking decision patterns and interaction structure without monitoring the contents of individual conversations.
- Capture content by reference: write full prompts and tool results to an access-controlled object store keyed by
trace_id/span_id, and put only the reference on the span. This keeps spans small (large attributes bloat backends and hit size limits) and lets you apply a different retention and access policy to the content than to the trace. - Capture content inline only in non-production or for opted-in sessions, with redaction at the SDK or collector layer for known PII patterns.
Whichever you choose, decide it explicitly. The default failure mode we see is the opposite: someone enables full-content capture to debug one issue, and it stays on.
Tool spans: where most real failures live
In our experience reviewing agent incidents, a large share of "the model got it wrong" turns out to be "the tool gave the model something misleading." Tool spans deserve more attributes than model spans.
- Record the contract, not just the call. Was the argument JSON valid against the schema? Did the tool return an error the model could read? Did the model retry with changed arguments or identical ones?
- Distinguish empty from failed.
results=[]and a 500 are different events with different mitigations. The model often treats an empty result as a fact about the world. - Link retries. Use
gen_ai.tool.call.idand a customretry.ofattribute so a retry storm collapses into one visible pattern. - Mark side effects. An attribute like
tool.side_effect=writelets you filter to the runs that changed external state, which is the first thing you want during an incident.
If your tools are thin wrappers over internal APIs, the same discipline you apply to idempotency and error contracts on the API side matters here; our API development and integration work often starts with making those contracts explicit enough for an agent to use safely.
Multi-agent traces: propagate context across boundaries
When an orchestrator spawns sub-agents, each sub-agent must continue the same trace. In-process this is automatic with context propagation; across queues, HTTP calls or A2A/MCP boundaries you must inject and extract W3C traceparent explicitly. Two cautions:
- Fan-out can make traces enormous. A planner that spawns many parallel workers creates very wide traces. Consider span links (rather than parent-child) for fire-and-forget workers whose lifetime outlasts the parent run.
- Long-running runs break the "one trace per request" assumption. An agent that waits on human approval for hours should end one trace and start a linked one, with
gen_ai.conversation.idas the stable join key.
For the architecture side of this problem, see our article on multi-agent orchestration that survives contact with production.
Sampling: keep the runs you will need to read
Agent traces are large and mostly boring. Storing everything at full fidelity is expensive; head sampling (deciding at the start) throws away exactly the runs you need, because you cannot know at the start which run will fail.
Tail-based sampling decides after the trace completes. A practical policy for LLM applications, described in Trace Sampling for LLM Apps, keeps: all error traces, slow traces above your latency budget, expensive traces above a cost threshold, evaluation and canary traffic at 100%, and a small percentage of healthy traffic (that article suggests about 5%). The OpenTelemetry Collector's tail sampling processor supports status_code, latency, string_attribute, numeric_attribute, probabilistic, and and composite policies (per the processor's README).
processors:
tail_sampling:
decision_wait: 120s # must exceed your slowest agent run
num_traces: 200000 # memory-bound: budget for it
policies:
- name: errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: slow-runs
type: latency
latency: { threshold_ms: 60000 }
- name: expensive-runs
type: numeric_attribute
numeric_attribute: { key: agent.cost_usd, min_value: 1.0 }
- name: max-steps-or-guardrail
type: string_attribute
string_attribute: { key: agent.stop_reason, values: [max_steps, guardrail, budget] }
- name: canary
type: string_attribute
string_attribute: { key: deployment.environment, values: [canary, eval] }
- name: baseline
type: probabilistic
probabilistic: { sampling_percentage: 5 }Policies are evaluated with OR semantics: a trace is kept if any policy matches. Three operational constraints bite agent workloads in particular:
decision_waitversus run length. The default is 30 seconds. Agent runs routinely exceed that. If the wait is shorter than the run, the processor decides on an incomplete trace. Set it above your slowest expected run and budget memory accordingly, sincenum_tracesbounds in-memory traces.- Same-collector routing. The processor requires all spans of a trace to reach the same collector instance, so you need a trace-ID-aware load-balancing tier in front of it.
agent.cost_usdis yours to compute. Token counts are standard; price is not. Compute cost at the SDK or in a processor from token counts and a pricing table, and remember it will drift when prices change.
Sampling strategy also interacts with metrics: compute rates and percentiles from metrics emitted before sampling, never from sampled traces.
Failure modes in the tracing system itself
| Failure mode | Symptom | Mitigation |
|---|---|---|
| Content capture left on | PII or secrets in the trace backend | Default-off, config-as-code, scan spans in CI, separate retention for content store |
| Oversized span attributes | Dropped attributes, rejected exports | Store content by reference; cap attribute length |
decision_wait too short | Partial traces, missing root span | Set above p99 run duration; alert on traces without a root |
| Context lost at async boundary | Orphan sub-agent traces | Explicit traceparent propagation; test it |
| Attribute dialect drift | Dashboards go blank after a library upgrade | Pin instrumentation versions; normalize in collector |
| Success-only instrumentation | Silent failures invisible | Map agent-level failure conditions to span status |
| High-cardinality attributes on metrics | Metrics bill explodes | Keep IDs on spans, not on metric labels |
A debugging workflow the traces should support
Design backward from the questions you will ask at 2 a.m.
- Find the cohort. Filter by
agent.config_hash,agent.stop_reasonand time. Is it one version, one tool, one tenant? - Open a representative trace and read the step list top to bottom: decision per step, tool results, finish reasons.
- Find the first bad step, not the last. Look for the first empty/malformed tool result or the first truncated context.
- Reproduce offline. With content captured by reference, replay the exact messages and tool results against the same config hash.
- Turn it into a regression case. Add the trace to your eval set so the fix is tested forever.
If your trace UI cannot support steps 1 to 3 in under a few minutes, the instrumentation is incomplete regardless of how many attributes you emit.
Decision checklist
- Does every run have one root
invoke_agentspan with a stableconversation.idandconfig_hash? - Are step boundaries explicit, with the decision recorded?
- Do agent-level failures (step budget, guardrail, budget cap) set span status?
- Are tool spans recording argument validity, result size, empty-vs-error, retries and side effects?
- Is prompt/tool content captured deliberately (reference store, retention, access control), not accidentally?
- Does tail sampling keep errors, slow runs, expensive runs, canaries and a baseline, with
decision_waitlonger than your slowest run? - Is trace context propagated across every async and cross-service boundary?
- Is there a path from a production trace to an eval case?
- Is prompt injection visible: do you tag spans where untrusted content entered the context? (See our piece on defending against prompt injection in LLM applications.)
Conclusion
Agent observability is a design problem before it is a tooling problem. Mirror the agent loop in your spans, turn agent-level failures into span status, record the decisions and tool contracts that explain behavior, capture content deliberately, and sample in a way that preserves the runs you will need to read. Do that, and the next silent failure becomes a query instead of a mystery.
Sources
- How we built our multi-agent research system, Anthropic Engineering
- Inside the LLM Call: GenAI Observability with OpenTelemetry
- OpenTelemetry GenAI Semantic Conventions Explained, Dash0
- OpenTelemetry GenAI semantic conventions repository
- Tail Sampling Processor, OpenTelemetry Collector Contrib
- Trace Sampling for LLM Apps: Keep the Spans That Matter, Drop the Rest
Working on this?
Syslabs' engineering team does architecture reviews on agent observability, evals and production LLM systems. Book a 30-minute architecture review
Internal Links
- LLM eval pipeline
- multi-agent orchestration
- prompt injection
- API development and integration
- custom software development