TL;DR: An agent failure is rarely a failed request. It is a bad decision made three steps earlier that only surfaced as a wrong answer. Traces that debug these failures need a span hierarchy that mirrors the agent loop (run, step, model call, tool call), attributes that explain decisions rather than just timing them, a deliberate policy for capturing prompt content, and tail-based sampling that keeps the interesting runs. Standard request tracing gives you none of this by default.

Why classic tracing falls short for agents

A conventional service trace answers "which hop was slow or errored?" That works because the call graph is fixed by code. An agent's call graph is chosen at runtime by a model. Two runs with the same input can take different paths, call different tools, and finish in different step counts. Anthropic's engineering team put it plainly when describing their multi-agent research system: agents "make dynamic decisions and are non-deterministic between runs, even with identical prompts," which makes debugging harder, and full production tracing is what let them diagnose failures systematically (How we built our multi-agent research system).

Three properties make agent failures different from ordinary service failures:

  1. Silent failures. The run returns HTTP 200 with a fluent, wrong answer. No span has status=ERROR. Your error-rate dashboard stays green.
  2. Causal distance. The root cause (a tool returned an empty list, the model treated it as "no results exist") sits several steps before the visible symptom.
  3. Compounding. A small mistake in step 2 changes the context for every later step. The same source article notes that minor system failures can be catastrophic for agents because errors compound.

So the design question is not "how do we log LLM calls?" It is "what would we need to see to reconstruct why the agent did what it did?"

The span hierarchy: mirror the loop, not the framework

Most agent runtimes are a loop: build context, call the model, parse a decision, maybe execute tools, repeat until a stop condition. Your trace should have the same shape.

text
invoke_agent  (the whole run; one per user request)
├── plan / step 1
│   ├── chat            (model call: decision + token usage)
│   ├── execute_tool    (search_orders)
│   └── execute_tool    (get_customer)
├── step 2
│   ├── retrieval       (vector lookup, if you do RAG)
│   └── chat
├── step 3
│   └── chat            (final answer; finish_reason=stop)
└── (guardrail / validation spans as siblings)

The OpenTelemetry GenAI semantic conventions define roughly this vocabulary. Operation names include invoke_agent, invoke_workflow, plan, chat, execute_tool, retrieval and memory operations (Dash0's walkthrough of the conventions). The OpenTelemetry project's own GenAI write-up uses the same three-level shape: an invoke_agent root, chat children for each LLM call, and execute_tool spans for tool invocations (Inside the LLM Call: GenAI Observability with OpenTelemetry).

A "step" span is worth adding

The conventions do not require an explicit per-iteration span, but we recommend one as a custom grouping span. Without it, a run with 14 model calls and 22 tool calls is a flat list of 36 children. With it, "the agent looped 9 times" is visible at a glance, and you can attach a step index and the stop/continue decision to one place.

Pseudo-code for the loop

python
with tracer.start_as_current_span("invoke_agent", attributes={
    "gen_ai.operation.name": "invoke_agent",
    "gen_ai.agent.name": "support-triage",
    "gen_ai.agent.version": AGENT_VERSION,      # prompt+tool-set hash, see below
    "gen_ai.conversation.id": conversation_id,
}) as run:
    for step in range(MAX_STEPS):
        with tracer.start_as_current_span("agent.step", attributes={"agent.step.index": step}) as s:
            resp = call_model(messages)           # emits a `chat` span
            s.set_attribute("agent.step.decision", classify(resp))  # tool_call | final | clarify
            if resp.tool_calls:
                for tc in resp.tool_calls:
                    run_tool(tc)                  # emits `execute_tool` span
                continue
            break
    else:
        run.set_attribute("agent.stop_reason", "max_steps")
        run.set_status(Status(StatusCode.ERROR, "step budget exhausted"))

Note the last two lines. A run that hits its step budget is a failure even though no exception was thrown. Convert agent-level failure conditions into span status, otherwise tail sampling and alerting cannot see them.

Attributes that explain decisions

Timing and token counts are table stakes. The attributes that let you debug are the ones that record what the agent knew and chose.

Question during an incidentAttribute / signal you need
Which version of the agent was this?gen_ai.agent.version, plus a hash of system prompt + tool schemas + model ID
Did the model hit a limit?gen_ai.response.finish_reasons (length vs stop vs tool_calls)
What did it cost?gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, cache-read tokens
Which tool, with what call identity?gen_ai.tool.name, gen_ai.tool.call.id
Why did the loop end?custom agent.stop_reason: final_answer, max_steps, budget, guardrail, tool_error
Was the context truncated or summarized?custom agent.context.tokens, agent.context.truncated
Was a tool result empty or malformed?custom tool.result.size, tool.result.empty, tool.schema_valid
Which conversation does this belong to?gen_ai.conversation.id

The gen_ai.* names above come from the OpenTelemetry GenAI conventions; the agent.* and tool.* names are our own, and you should namespace your own attributes so they cannot collide with future standard ones.

Stability caveat: pin and centralize attribute names

The GenAI attributes are still at "Development" stability, and the conventions have moved to a dedicated repository for faster iteration (per the same Dash0 walkthrough). Different instrumentation libraries emit different dialects. Two practical mitigations: define attribute-name constants in one module of your own code, and normalize dialects at the collector (for example with the transform processor) rather than in dashboards.

Versioning is the highest-leverage attribute

When quality regresses on Tuesday, the first question is "what changed?" An agent's behavior is a function of model, system prompt, tool descriptions, retrieval index and temperature. Hash the prompt text and the serialized tool schemas into a single agent.config_hash, and stamp it on the root span. Then "failures started when hash a41f… rolled out" is a query, not an investigation. This also gives your LLM eval pipeline a join key between offline eval runs and production traces. If you are building that side, see our write-up on how to build an LLM eval pipeline from golden datasets to CI gates.

Content capture: the privacy and cost decision

Prompts, tool arguments and completions are what you actually read when debugging, and they are also where personal data, secrets and customer content live. The OpenTelemetry conventions treat this deliberately: prompt and completion content is excluded from default span attributes and exposed through opt-in attributes/events instead. As the OpenTelemetry post puts it for Copilot's instrumentation, by default no prompt content or tool arguments are captured because they can contain sensitive data (same OpenTelemetry post). Opt-in mechanics differ by library, so check yours.

A workable policy has three tiers:

  1. Always capture structure, never content: step counts, tool names, argument shapes (keys and types), result sizes, token counts, hashes, error types. This is cheap and safe, and it answers most "what path did the agent take?" questions. Anthropic describes taking this approach to production monitoring: tracking decision patterns and interaction structure without monitoring the contents of individual conversations.
  2. Capture content by reference: write full prompts and tool results to an access-controlled object store keyed by trace_id/span_id, and put only the reference on the span. This keeps spans small (large attributes bloat backends and hit size limits) and lets you apply a different retention and access policy to the content than to the trace.
  3. Capture content inline only in non-production or for opted-in sessions, with redaction at the SDK or collector layer for known PII patterns.

Whichever you choose, decide it explicitly. The default failure mode we see is the opposite: someone enables full-content capture to debug one issue, and it stays on.

Tool spans: where most real failures live

In our experience reviewing agent incidents, a large share of "the model got it wrong" turns out to be "the tool gave the model something misleading." Tool spans deserve more attributes than model spans.

  • Record the contract, not just the call. Was the argument JSON valid against the schema? Did the tool return an error the model could read? Did the model retry with changed arguments or identical ones?
  • Distinguish empty from failed. results=[] and a 500 are different events with different mitigations. The model often treats an empty result as a fact about the world.
  • Link retries. Use gen_ai.tool.call.id and a custom retry.of attribute so a retry storm collapses into one visible pattern.
  • Mark side effects. An attribute like tool.side_effect=write lets you filter to the runs that changed external state, which is the first thing you want during an incident.

If your tools are thin wrappers over internal APIs, the same discipline you apply to idempotency and error contracts on the API side matters here; our API development and integration work often starts with making those contracts explicit enough for an agent to use safely.

Multi-agent traces: propagate context across boundaries

When an orchestrator spawns sub-agents, each sub-agent must continue the same trace. In-process this is automatic with context propagation; across queues, HTTP calls or A2A/MCP boundaries you must inject and extract W3C traceparent explicitly. Two cautions:

  • Fan-out can make traces enormous. A planner that spawns many parallel workers creates very wide traces. Consider span links (rather than parent-child) for fire-and-forget workers whose lifetime outlasts the parent run.
  • Long-running runs break the "one trace per request" assumption. An agent that waits on human approval for hours should end one trace and start a linked one, with gen_ai.conversation.id as the stable join key.

For the architecture side of this problem, see our article on multi-agent orchestration that survives contact with production.

Sampling: keep the runs you will need to read

Agent traces are large and mostly boring. Storing everything at full fidelity is expensive; head sampling (deciding at the start) throws away exactly the runs you need, because you cannot know at the start which run will fail.

Tail-based sampling decides after the trace completes. A practical policy for LLM applications, described in Trace Sampling for LLM Apps, keeps: all error traces, slow traces above your latency budget, expensive traces above a cost threshold, evaluation and canary traffic at 100%, and a small percentage of healthy traffic (that article suggests about 5%). The OpenTelemetry Collector's tail sampling processor supports status_code, latency, string_attribute, numeric_attribute, probabilistic, and and composite policies (per the processor's README).

yaml
processors:
  tail_sampling:
    decision_wait: 120s        # must exceed your slowest agent run
    num_traces: 200000         # memory-bound: budget for it
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow-runs
        type: latency
        latency: { threshold_ms: 60000 }
      - name: expensive-runs
        type: numeric_attribute
        numeric_attribute: { key: agent.cost_usd, min_value: 1.0 }
      - name: max-steps-or-guardrail
        type: string_attribute
        string_attribute: { key: agent.stop_reason, values: [max_steps, guardrail, budget] }
      - name: canary
        type: string_attribute
        string_attribute: { key: deployment.environment, values: [canary, eval] }
      - name: baseline
        type: probabilistic
        probabilistic: { sampling_percentage: 5 }

Policies are evaluated with OR semantics: a trace is kept if any policy matches. Three operational constraints bite agent workloads in particular:

  • decision_wait versus run length. The default is 30 seconds. Agent runs routinely exceed that. If the wait is shorter than the run, the processor decides on an incomplete trace. Set it above your slowest expected run and budget memory accordingly, since num_traces bounds in-memory traces.
  • Same-collector routing. The processor requires all spans of a trace to reach the same collector instance, so you need a trace-ID-aware load-balancing tier in front of it.
  • agent.cost_usd is yours to compute. Token counts are standard; price is not. Compute cost at the SDK or in a processor from token counts and a pricing table, and remember it will drift when prices change.

Sampling strategy also interacts with metrics: compute rates and percentiles from metrics emitted before sampling, never from sampled traces.

Failure modes in the tracing system itself

Failure modeSymptomMitigation
Content capture left onPII or secrets in the trace backendDefault-off, config-as-code, scan spans in CI, separate retention for content store
Oversized span attributesDropped attributes, rejected exportsStore content by reference; cap attribute length
decision_wait too shortPartial traces, missing root spanSet above p99 run duration; alert on traces without a root
Context lost at async boundaryOrphan sub-agent tracesExplicit traceparent propagation; test it
Attribute dialect driftDashboards go blank after a library upgradePin instrumentation versions; normalize in collector
Success-only instrumentationSilent failures invisibleMap agent-level failure conditions to span status
High-cardinality attributes on metricsMetrics bill explodesKeep IDs on spans, not on metric labels

A debugging workflow the traces should support

Design backward from the questions you will ask at 2 a.m.

  1. Find the cohort. Filter by agent.config_hash, agent.stop_reason and time. Is it one version, one tool, one tenant?
  2. Open a representative trace and read the step list top to bottom: decision per step, tool results, finish reasons.
  3. Find the first bad step, not the last. Look for the first empty/malformed tool result or the first truncated context.
  4. Reproduce offline. With content captured by reference, replay the exact messages and tool results against the same config hash.
  5. Turn it into a regression case. Add the trace to your eval set so the fix is tested forever.

If your trace UI cannot support steps 1 to 3 in under a few minutes, the instrumentation is incomplete regardless of how many attributes you emit.

Decision checklist

  • Does every run have one root invoke_agent span with a stable conversation.id and config_hash?
  • Are step boundaries explicit, with the decision recorded?
  • Do agent-level failures (step budget, guardrail, budget cap) set span status?
  • Are tool spans recording argument validity, result size, empty-vs-error, retries and side effects?
  • Is prompt/tool content captured deliberately (reference store, retention, access control), not accidentally?
  • Does tail sampling keep errors, slow runs, expensive runs, canaries and a baseline, with decision_wait longer than your slowest run?
  • Is trace context propagated across every async and cross-service boundary?
  • Is there a path from a production trace to an eval case?
  • Is prompt injection visible: do you tag spans where untrusted content entered the context? (See our piece on defending against prompt injection in LLM applications.)

Conclusion

Agent observability is a design problem before it is a tooling problem. Mirror the agent loop in your spans, turn agent-level failures into span status, record the decisions and tool contracts that explain behavior, capture content deliberately, and sample in a way that preserves the runs you will need to read. Do that, and the next silent failure becomes a query instead of a mystery.

Sources

Working on this?

Syslabs' engineering team does architecture reviews on agent observability, evals and production LLM systems. Book a 30-minute architecture review

  • LLM eval pipeline
  • multi-agent orchestration
  • prompt injection
  • API development and integration
  • custom software development