Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIAI TestingDevOps

AI Agent Tracing: Spans, Tool Calls and OpenTelemetry

AI agent tracing explained: what an agent trace is made of, how OpenTelemetry names its spans, how context crosses MCP, and where a trace stops being evidence.

Published on:

OVERVIEW

Five frontier models flagged failed agent runs almost every time and almost never found the step where the failure began, when reading traces reduced to metadata or to fields mapped onto the OpenTelemetry and OpenInference conventions. TelemetrySuffBench, a controlled benchmark published as a preprint in August 2026, built those views itself from 312 constructed traces: detection F1 stayed at 99.5% to 100% and origin-step accuracy at or below 0.5%, against 33.8% to 97.2% with full telemetry. AI agent tracing records one agent run as a tree of spans, one for each agent invocation, model call and tool call.

What you put in those spans decides which questions the trace can answer. This guide traces one scripted multi-step agent run (stub model, tools and warehouse) with the OpenTelemetry JavaScript SDK: its 16 spans pinpointed a failed tool call and its retry, and could not show that the two calls left two records behind.

Overview

AI agent tracing is the practice of recording each agent run as a trace: a tree of timed spans for every agent invocation, model call, tool call and retrieval step, linked by one trace ID and parent span IDs. A trace shows the path the agent took, including handoffs, how long each step ran and what the caller saw.

What are the key concepts in tracing AI agents?

  • Span: A span is one unit of work in an agent run, such as a model call or a tool call. It carries a name, a start and end time, a status, attributes and the ID of its parent span.
  • Trace ID: A trace ID is a 32-character hex identifier shared by every span in one agent run. The W3C traceparent value carries the trace ID across service boundaries, together with the 16-character ID of the calling span.
  • OpenTelemetry GenAI conventions: The OpenTelemetry GenAI semantic conventions name the spans in an agent trace, including invoke_agent, chat and execute_tool. As of 5 October 2026 they are marked Development and kept in a dedicated repository with no tagged release.
  • Context propagation: Context propagation carries the trace ID from one agent or tool to the next, in HTTP headers and, for MCP, in the request's params._meta field. When a caller does not inject it, one user request splits into separate traces.
  • Opt-in content: Under the OpenTelemetry GenAI conventions, prompts, model outputs, tool arguments and tool results are opt-in attributes that instrumentations should not capture by default. Decide per tool what an agent trace may hold before you turn them on.

Does a trace prove what an AI agent did?

A trace shows the calls the agent's stack recorded and the status each caller saw. It does not show the state of the system a tool wrote to, so a write that timed out may still have landed. TestMu AI's Agent Assurance checks that effect against staging before release, using changed files, artifacts and read-only record checks you provide, and reports what it could not verify separately, outside the pass rate.

What Is AI Agent Tracing?

AI agent tracing is the practice of recording every step of an agent run as a span and joining those spans into one trace, so you can see afterwards which model calls, tool calls, retrievals and handoffs the agent made, in what order, how long each took and what the caller got back.

The vocabulary comes from OpenTelemetry, whose traces documentation describes a trace as the path of a request through your application and says a span "represents a unit of work or operation". Agent tracing, also called agentic tracing, keeps that model and adds names for the steps only agents have.

  • Agent trace - the trace of one agent run: every span that shares its trace ID, under a root span that has no parent. Tracing is the practice of producing it.
  • Conversation - several traces grouped by a conversation ID, such as a thread or session ID, when one exchange spans many runs.

Spans and IDs work exactly as they do in distributed tracing. What changes is who picks the path: a model chooses each next step at run time, so two runs of the same request can produce two different trees.

Tracing is one signal inside AI agent observability, which also covers metrics, evaluations and the tools that store all three.

What Is an Agent Trace Made Of?

An agent trace is made of spans: a workflow span when several agents are coordinated, an invoke_agent span for each agent, a model call span such as chat, an execute_tool span for each tool call, and retrieval, memory and MCP spans where the run uses them. The names and kinds below come from the OpenTelemetry GenAI agent span conventions and the model span and MCP pages beside them, as they read on 5 October 2026.

SpanWhat it recordsName under the conventionsKind
Workflow runA coordinated process made of several agents or other GenAI operations. Not reported for a standalone agent.invoke_workflow {gen_ai.workflow.name}INTERNAL
Agent, in processOne agent running inside your own process. The conventions name LangChain and CrewAI agents as examples.invoke_agent {gen_ai.agent.name}INTERNAL
Agent, remoteA call to an agent that runs behind a remote service, such as a hosted agent API.invoke_agent {gen_ai.agent.name}CLIENT
PlanningA planning or task decomposition phase, reported only when the instrumentation can tell it apart from ordinary inference.plan {gen_ai.agent.name}INTERNAL
Model callOne request to a model, with the provider, the model and the token counts.{gen_ai.operation.name} {gen_ai.request.model}, such as chat followed by the model nameCLIENT, or INTERNAL in process
Tool callOne tool execution. The tool name is required, and arguments and results are opt-in.execute_tool {gen_ai.tool.name}INTERNAL
RetrievalA request that fetches context from a vector database or search system.retrieval {gen_ai.data_source.id}CLIENT
Memory operationCreating, searching, updating or deleting memory records, or creating and deleting a memory store.The operation name alone, such as search_memoryCLIENT, or INTERNAL in process
MCP tool callA tools/call request to an MCP server, traced on both sides.tools/call {gen_ai.tool.name}CLIENT on the caller, SERVER on the MCP server
HandoffControl passing from one agent to another. It usually shows up as another invoke_agent span or as a tool call, unless the instrumentation adds a custom operation name.None. The list of 18 well-known operation names has no handoff entry.Follows the span that carries it

Kind follows where the work runs: INTERNAL for an operation that stays inside your process and CLIENT for an outgoing remote call.

Worked Example: One Run, Sixteen Spans

To show those rows in a trace, a Node script written for this guide emits spans for one stock-transfer request and prints them as a tree. An inventory-desk agent checks stock, creates a transfer (the first attempt times out and the model calls the tool again), then hands the pickup to a carrier-agent behind an MCP-style tools/call message.

Everything below the span API is a stub: the model, the tools, the MCP server and the warehouse are sample code with sample data, and the latencies are timers. The script uses two tracer providers, one per service, that share an in-memory exporter. It was run on 5 October 2026 with Node 25.5.0, @opentelemetry/api 1.9.1 and @opentelemetry/sdk-trace-base 2.11.0 on Windows 11.

node trace-agent-run.mjs

This is the console output of that run, unedited up to its summary line:

node v25.5.0 | @opentelemetry/api 1.9.1 | sdk-trace-base 2.11.0 | core 2.11.0
propagation: traceparent injected into params._meta | content capture: off (default)

MCP request that crossed the process boundary:
{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"book_pickup","arguments":{"transfer_id":"TR-20931"},"_meta":{"traceparent":"00-7499b447f30d8ae5a2dd6e4ab34e43c5-c904d52a6339563b-01"}}}

traces produced by one user request: 1

trace_id 7499b447f30d8ae5a2dd6e4ab34e43c5   16 spans   2 service(s): inventory-desk, carrier-service
span_id           parent_id         kind      status  duration   span
a22e96f2f56194d4  (root)            INTERNAL  UNSET    355.8 ms  invoke_workflow stock_transfer
9f2a2b1da2b7091d  a22e96f2f56194d4  INTERNAL  UNSET    355.3 ms    invoke_agent inventory-desk
d01c9ddbe4d6812c  9f2a2b1da2b7091d  CLIENT    UNSET     43.8 ms      chat sample-model  [finish=tool_calls]
d12ab1cb66a3aca6  9f2a2b1da2b7091d  INTERNAL  UNSET     14.9 ms      execute_tool check_stock  [call.id=call_1]
de88eed2d8bcb8c1  9f2a2b1da2b7091d  CLIENT    UNSET     32.7 ms      chat sample-model  [finish=tool_calls]
8a638318e3283763  9f2a2b1da2b7091d  INTERNAL  ERROR     47.5 ms      execute_tool create_transfer  [call.id=call_2, error.type=timeout]
4d5f0c48a33a641a  9f2a2b1da2b7091d  CLIENT    UNSET     31.2 ms      chat sample-model  [finish=tool_calls]
4e1c0540001c33da  9f2a2b1da2b7091d  INTERNAL  UNSET     31.6 ms      execute_tool create_transfer  [call.id=call_3]
eb9c4f4bb79ea0b8  9f2a2b1da2b7091d  CLIENT    UNSET     31.4 ms      chat sample-model  [finish=tool_calls]
c904d52a6339563b  9f2a2b1da2b7091d  CLIENT    UNSET     79.4 ms      tools/call book_pickup
3e193128e8c7376b  c904d52a6339563b  SERVER    UNSET     78.2 ms        tools/call book_pickup
170e978312164bdb  3e193128e8c7376b  INTERNAL  UNSET     78.1 ms          invoke_agent carrier-agent
edb30dd15984674c  170e978312164bdb  CLIENT    UNSET     30.4 ms            chat sample-model  [finish=tool_calls]
427095ec4c7944af  170e978312164bdb  INTERNAL  UNSET     15.4 ms            execute_tool reserve_dock_slot  [call.id=call_r1]
ae4b354001d6c9c7  170e978312164bdb  CLIENT    UNSET     31.9 ms            chat sample-model  [finish=stop]
ad8fca044107b5da  9f2a2b1da2b7091d  CLIENT    UNSET     40.8 ms      chat sample-model  [finish=stop]

summary: 16 spans | 7 model calls | 5 tool calls, 1 with status ERROR | 3570 input + 233 output tokens
  • Trace ID - all 16 spans carry the trace ID that starts 7499b447: 11 come from inventory-desk and 5 from carrier-service.
  • Parent span ID - the root span has no parent, and invoke_agent inventory-desk lists the root's span ID, a22e96f2f56194d4, as its parent.
  • Retry - create_transfer appears twice: call_2 with status ERROR and error.type=timeout, then call_3. The conventions say a request retried automatically stays in one span that covers "the duration of the logical operation with all retries", so in a trace that follows the conventions, two spans mean the model decided to call the tool again.
  • Unset status - UNSET on the other 15 spans is not a warning: OpenTelemetry's default span status is Unset, which means the operation completed without an error.
  • Handoff - no span of its own: it appears as the tools/call book_pickup client span, the matching server span in the other service, and the invoke_agent carrier-agent span under it.

Read one of your own traces the same way: each tool call should appear once on the caller's side, and every span except the root should name a parent that is in the trace.

How Does OpenTelemetry Name Agent Spans Today?

OpenTelemetry for AI agents names each span by its operation: invoke_workflow for a coordinated run, invoke_agent for one agent, a model operation such as chat, and execute_tool for a tool call. Those names are defined in the GenAI semantic conventions, and as of 5 October 2026 they can still change:

  • Status - the agent spans, model spans, events and MCP pages are each marked Development, and so is every GenAI span and attribute in the table above. The stable entries are generic ones such as error.type.
  • Location - the opentelemetry.io page for these conventions is now titled "Moved: Generative AI semantic conventions" and points to a dedicated OpenTelemetry repository created in May 2026.
  • Releases - that repository has no tagged release, and the agent spans page last changed on 30 September 2026. Record the commit or date your instrumentation follows.

Before you add attributes of your own, apply the conventions' rules on identity and duplication:

  • Conversation ID - set gen_ai.conversation.id only when a real one exists. A new UUID, a trace identifier or a hash of the request content should not be used as a fallback.
  • Agent ID - gen_ai.agent.id is defined as the stable identifier of a hosted agent resource, and recording in-memory instance IDs on it is not recommended. An in-process agent span carries the agent's name, so the script puts the release on the resource as service.version.
  • One span per tool call - a call that is both a generic and a specialized tool call, such as an Agent Skill exposed as a tool, should not be recorded as two different spans. An MCP instrumentation that can tell the call is already traced should add its attributes to the existing span.

OpenTelemetry's names are one of several vocabularies you will meet. The OpenAI Agents SDK for Python traces by default with span types of its own, listed in its tracing documentation, and the OpenInference specification tags every span with a required openinference.span.kind attribute:

What happenedOpenTelemetry GenAI operationOpenInference span kindOpenAI Agents SDK default span
A run that coordinates several agentsinvoke_workflowCHAIN, its kind for a starting point or a link between stepstrace() around the run and task_span() for each runner invocation
One agent workinginvoke_agentAGENTagent_span()
A model callchat, generate_content or text_completionLLMgeneration_span(), plus a turn_span() for each model turn
A tool callexecute_toolTOOLfunction_span()
A handoffNo dedicated operationNo dedicated kindhandoff_span()
A guardrail checkNo dedicated operationGUARDRAILguardrail_span()

The gaps matter when you query across sources: a dashboard that counts handoffs by span name works for one of these vocabularies and may find nothing in the other two. Which backend stores the spans is a separate decision, covered in the comparison of AI observability tools. The roundup of AI agent observability tools compares what each one shows for a multi-step run, including handoffs, sessions and MCP calls.

How Does Trace Context Cross Agents, Tools and MCP Servers?

Multi agent tracing holds together only when every hop passes two identifiers to the next process: the trace ID and the caller's span ID. Both travel in the traceparent value defined by the W3C Trace Context recommendation. This is the value the script sent, 00-7499b447f30d8ae5a2dd6e4ab34e43c5-c904d52a6339563b-01, field by field:

  • Version - 00.
  • Trace ID - 7499b447f30d8ae5a2dd6e4ab34e43c5, which is 32 hex characters or 16 bytes.
  • Parent ID - c904d52a6339563b, which is 16 hex characters or 8 bytes. It is the span ID of the tools/call book_pickup client span in the output above.
  • Trace flags - 01. The last bit set means the caller may have recorded trace data.

The value has to be sent only when a hop leaves the process:

  • Handoff inside one process - nothing to send: the SDK's active context makes the next agent's span a child of the current one.
  • Call to an agent in another service - the SDK's propagator writes traceparent into the HTTP request headers and the receiving service reads it.
  • Handoff through a queue - the receiving run may start much later, so OpenTelemetry's traces documentation suggests a span link, which associates the later trace with the span that queued the work.

An MCP tool call needs more than headers. OpenTelemetry's MCP semantic conventions explain that HTTP trace context propagation "only covers the HTTP request, but not the individual messages client and server exchange within the request/response streams". Instrumentations should therefore inject the context into the MCP request's params._meta property, with the keys written unprefixed, and the receiver uses it as the remote parent.

MCP documents the same carrier from its side. SEP-414, a Final standards-track proposal whose rule is now in the MCP specification (2026-07-28 revision), warns that a namespaced key such as io.modelcontextprotocol.traceparent "will break traces and log correlation". These are the two functions from the script, where context, propagation and ROOT_CONTEXT come from @opentelemetry/api and PROPAGATE is the script's own flag:

// MCP client side. The tools/call CLIENT span is the active span when this runs.
function buildToolCall(id, name, args) {
  const message = { jsonrpc: "2.0", id, method: "tools/call", params: { name, arguments: args, _meta: {} } };
  if (PROPAGATE) propagation.inject(context.active(), message.params._meta); // writes traceparent
  return message;
}

// MCP server side. The caller's context comes out of params._meta and becomes the remote parent.
function remoteParentOf(message) {
  return propagation.extract(ROOT_CONTEXT, message.params._meta ?? {});
}

With the inject call in place, the server span in the first output lists the client span as its parent. The script was then run with --no-propagate, which skips the inject call, and the request left with an empty _meta:

MCP request that crossed the process boundary:
{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"book_pickup","arguments":{"transfer_id":"TR-20931"},"_meta":{}}}

traces produced by one user request: 2

The carrier service's five spans landed in a trace of their own, rooted at the server span:

trace_id 190dc34b4f35e46a41fcdc2f576f98b3   5 spans   1 service(s): carrier-service
span_id           parent_id         kind      status  duration   span
dfe8201663573a61  (root)            SERVER    UNSET     87.9 ms  tools/call book_pickup
ab673b7be6395034  dfe8201663573a61  INTERNAL  UNSET     87.8 ms    invoke_agent carrier-agent
327068fe11698f8a  ab673b7be6395034  CLIENT    UNSET     30.7 ms      chat sample-model  [finish=tool_calls]
15e4406da8644759  ab673b7be6395034  INTERNAL  UNSET     25.1 ms      execute_tool reserve_dock_slot  [call.id=call_r1]
d2fb9adeda8ae540  ab673b7be6395034  CLIENT    UNSET     31.4 ms      chat sample-model  [finish=stop]

Nothing in either trace flags the break. The only symptom is a tools/call server span sitting at the root of its own trace, which is something you can search a backend for. A caller you do not trace at all, such as a third-party MCP client, leaves the same shape.

Some hops cannot carry context from your side. Microsoft Agent Framework's observability documentation says, for Python, that its params._meta injection covers MCP sessions the agent process opens and does not apply to hosted or provider-managed MCP tools, where the provider's service runtime issues the tools/call message. If you need tracing through to the MCP server, its advice is a client-opened MCP transport.

What to assert at a handoff once the trace shows it, such as the fields the receiving agent needs, is covered in the guide to agent handoff testing.

What Should an Agent Trace Capture, and What Should It Redact?

Start from the default in the conventions, which is content off. The model span page of the OpenTelemetry GenAI conventions says instrumentations should not capture instructions, inputs and outputs by default and should offer an opt-in, and it documents these usage patterns:

  • Record nothing - the default.
  • Record on spans - suited to manageable volume where privacy regulations do not apply or the telemetry storage meets them, for example in pre-production environments.
  • Store externally - content in separate storage with references on the spans, recommended for production where telemetry volume is a concern or sensitive data needs to be handled securely.

Your instrumentation may not follow that default, so check its setting before the first production run:

  • OpenAI Agents SDK for Python - its documentation states that trace_include_sensitive_data is True by default, so generation and function spans store inputs and outputs.
  • Microsoft Agent Framework for Python - ENABLE_SENSITIVE_DATA defaults to false, and prompts, responses, function call arguments and results are logged only when you set it to true.

OpenTelemetry's guide to handling sensitive data puts the review on you: you are responsible for "understanding and reviewing the telemetry data emitted by any instrumentation libraries you use".

Redacting too much also has a cost: in TelemetrySuffBench, removing only the decision content reduced origin-step accuracy to zero for every model. The benchmark is synthetic, built on one family of injected fault, and its OpenTelemetry-compatible view is a conservative field set its authors chose, so read its figures as a controlled result. A workable starting rule is to keep identifiers and the link between a decision and the call it produced, and to mask the values.

In this run, opt-in content added one fact the default output lacked. With --capture-content, the two create_transfer spans show that the retry carried the same arguments as the call that timed out:

eb3b5ee89278999f  0274c4b1cb735614  INTERNAL  ERROR     47.8 ms      execute_tool create_transfer  [call.id=call_2, error.type=timeout, args={"sku":"SKU-7731","qty":40,"from":"PNQ-1","to":"BLR-2"}]
cc6a7d6a8e993fa1  0274c4b1cb735614  CLIENT    UNSET     31.3 ms      chat sample-model  [finish=tool_calls]
216692c76460c997  0274c4b1cb735614  INTERNAL  UNSET     31.7 ms      execute_tool create_transfer  [call.id=call_3, args={"sku":"SKU-7731","qty":40,"from":"PNQ-1","to":"BLR-2"}, result={"status":"accepted","transfer_id":"TR-20931"}]

A per-attribute policy makes that a decision you take once per tool:

What the span holdsDefault under the conventionsRiskPolicy to start from
Span name, status, operation, agent and tool names, tool call ID, model, token countsRequired, conditionally required or recommended, so recorded without opting inInternal names become visible to everyone with trace accessKeep. Samplers and cost reports read them
Tool arguments and resultsOpt-inAccount numbers, personal data and secrets passed as argumentsDecide per tool: keep identifiers, mask values
Prompts, outputs, system instructionsOpt-inPersonal data, and sizeOn spans before production, external storage with references in production
Retrieved documents and query textOpt-inDocument textRecord document IDs and scores, the shape the conventions' own example uses
traceparent and tracestateCarried between servicesThe W3C recommendation forbids personal or sensitive information in themIdentifiers only

Redact where the data is created when you can, and otherwise in the pipeline: the same OpenTelemetry guide lists the Collector's attribute, filter, redaction and transform processors. It also warns that hashing a user ID may not anonymize it, because hashes are reversible in practice when the input space is small and predictable. For the techniques themselves, see the guide to data masking.

How Do You Sample AI Agent Traces Without Losing the Failures?

Tracing AI agents with default settings keeps everything. The OpenTelemetry SDK environment variable specification lists parentbased_always_on as the default sampler and "no limit" as the default maximum size of a span attribute value, so every run is kept and a recorded prompt is stored at full length.

Compare your own volume with OpenTelemetry's thresholds before you drop anything. Its sampling documentation lists generating 1000 or more traces per second as a reason to sample and tens of small traces per second or lower as a reason not to. An agent in the second group pays for size per trace, so control size first:

  • Cap attribute size - set OTEL_SPAN_ATTRIBUTE_VALUE_LENGTH_LIMIT so one pasted document cannot become one enormous attribute, or move content to external storage as described above.
  • Weigh payloads first - paste a sample or redacted prompt or tool result into the JSON Size Analyzer to see how many bytes each field adds.

When volume does force sampling, the same sampling documentation describes head and tail sampling:

  • Head sampling - decides as early as possible, without inspecting the trace as a whole, so on its own it cannot ensure that every trace containing an error is kept.
  • Tail sampling - decides after considering all or most of the spans in a trace, so a policy can keep every trace that contains an error. It needs stateful components that can store a large amount of data.

Tail sampling has a default that long agent runs collide with. The OpenTelemetry Collector's tail sampling processor waits 30 seconds by default (decision_wait) and keeps 50,000 traces in memory (num_traces). Its README calls a span late when it arrives after its trace's sampling decision: late spans inherit that decision while it remains in the buffer, and after that, by default, they are buffered and judged as if the trace were new.

Read literally, those defaults mean a run that outlasts the wait is judged on the spans that arrived within it. An error three minutes in cannot rescue a trace that was already dropped, and at best survives as a fragment without the steps that led to it. Set decision_wait above your longest expected run, size num_traces to match, and add a status_code policy that keeps errors.

How Does AI Agent Tracing Feed Evaluation?

A trace is the input to evaluation at two levels: the whole run (did it reach the goal) and the single span (was this tool call the right one). A score attached to a span points at the step that went wrong.

The OpenTelemetry GenAI event conventions define the carrier: an event named gen_ai.evaluation.result that should be parented to the span being evaluated when possible, with these attributes:

  • gen_ai.evaluation.name - the name of the evaluation metric, and the only required attribute.
  • gen_ai.evaluation.score.value and gen_ai.evaluation.score.label - a number and a readable label, such as pass or fail.
  • gen_ai.evaluation.explanation - the evaluator's free-form explanation for the score.

Recorded this way, a score stays queryable beside the span it grades, whichever evaluator produced it. The guide to agent regression testing explains how to turn a failure that a scored trace exposes into a scenario that gates the next release.

Where Does a Trace Stop Being Evidence?

A trace is the record the agent's own stack emitted. That makes it strong evidence of what was called and what each caller saw, and weak evidence of the state those calls left behind. The run from the worked example ended with one more block, read straight from the stub warehouse:

warehouse ledger, read directly from the stub (no span records this):
transfer_id  sku       qty  from   to     written_by
TR-20930     SKU-7731   40  PNQ-1  BLR-2  call_2
TR-20931     SKU-7731   40  PNQ-1  BLR-2  call_3
transfers written for one request: 2

The trace showed one create_transfer that failed and one that succeeded. The warehouse holds two transfers for a request that asked for one, because the stub is written so the first call commits its row before the reply times out, the way a write behind a slow response can. None of the 16 spans records that.

A trace is also silent about work done outside your process, and a complete record still leaves the failing step hard to find:

  • Provider-side tools - the GenAI conventions' metric for tool calls per agent invocation counts only tools executed by the agent or framework and leaves out tools the model provider runs server-side, such as built-in web search or code execution. Your process never executes those, so expect the same gap in your tool spans.
  • Attribution - in a 2025 paper built on the Who&When dataset of failure logs from 127 LLM multi-agent systems, the best automated method identified the responsible agent with 53.5% accuracy and the decisive step with 14.2%.
QuestionCan the trace answer it?What answers it
Which tools ran, in what order, for how long?Yes, when each call has a spanThe trace
What arguments did the agent pass?Only when opt-in content is recordedThe trace with content on
Did the call return an error?Yes, from the span status and error.typeThe trace
Did the write land, and only once?No. A span shows what the caller sawA read-only query of the system the tool writes to
What did a remote agent or hosted connector do?Only when context was propagated and that side emits spansThat side's telemetry
Why did the agent choose this step?Only when decision content is recordedRecorded inputs and outputs for that step

An agent can also report a write that never happened, the opposite failure, which the guide to AI agent hallucination covers. In production, reconciling a sample of writes against the system of record is a job for AI agent monitoring.

Checking What a Trace Cannot Show With Agent Assurance

TestMu AI built Agent Assurance to test agents that act, before they are released. It reads your agent's code or spec, writes the scenarios, invokes the agent for real in staging and judges every acceptance criterion on evidence of what the run did. It is not an observability or tracing product, so it works beside your tracing backend, and a hook you write can pass it the tool calls that backend recorded.

The connection point is the profile: the reviewable hook scripts that tell Agent Assurance how to reach your agent. Its execute and collect hooks each write one JSON object, and the Rook CLI profiles and hooks guide defines the fields a tracing team will use:

  • Observed calls - the calls field holds the tool calls you observed, which lets a criterion assert that a tool was or was not called. A collect hook runs after a scenario closes, so it can read that scenario's spans from your tracing backend and return the calls it finds.
  • Extra fields - anything else the hook returns, such as a trace URL, is preserved as run evidence.
  • Empty calls list - return an empty list only when you observed that no calls occurred. Leaving the field out means calls were not observable, and criteria that need them come back Unable to Verify.
  • Late spans - a delay_seconds setting on the collect hook sets a minimum wait before collection, for spans that reach your backend after the agent replies.

For an agent like the stock-transfer example, the documented checks map onto the run this way:

  • Transfer record - whether staging holds exactly one transfer for the request can be settled only by a read-only check that you provide: a query tool on a stdio MCP server you approve, a read-only status endpoint or a collect hook that fetches the record. Without one, the criterion is Unable to Verify.
  • Both create_transfer calls - once the hook returns them, each call is compared with the agent's declared tools, and a criterion can name a tool the agent must never call.
  • Handoff to carrier-agent - the profile has to return the call that shows it, here book_pickup, before a criterion about the handoff can be checked.
  • Unable to Verify - the verdict for any criterion that no call, file, artifact or read-only check could settle. It counts neither as a pass nor as a failure, and it stays outside the pass rate.

Most eval and observability tools score what your agent said and recorded. Agent Assurance checks what the run changed, and reports what it could not verify:

AspectAgent AssuranceLLM observability
Tool callsAgainst declared toolsLogged, optionally scored
Side effectsFiles, artifacts, probesTrace data only
When it runsBefore release, in CIProduction, plus CI
Unverifiable resultsReported separatelyLeft unscored

The observability column describes the category's default approach, not any single product. Like eval and observability tools, Agent Assurance uses model judges, grades the reply as well as the action and runs in CI. What differs is what it accepts as proof of an action.

On a first run against an agent that records little, expect many criteria to come back Unable to Verify. The remedy is the tracing work above: record the tool calls, return them from the hook, and the unverified share falls.

Agent Assurance is pre-alpha and publicly installable, and it runs from the terminal as Rook CLI on macOS and Linux, and 64-bit Windows through npm or WSL. Install it from npm, the route that needs Node.js 22 or newer:

npm install -g @testmuai/rook

The Rook CLI install guide also covers Homebrew and a shell installer that verifies the checksum of what it downloads. Give every run a staging profile: Agent Assurance calls your agent as a user would, so each write the agent makes is real and the test cannot undo it.

Claude Code users can operate Rook CLI through the rook skill: install it with the first command below, open a new session, type / and select rook, then describe the profile you need in plain language:

npx @testmuai/rook-skill@latest install --agent claude-code
/rook Prepare a staging profile for the stock-transfer agent in this repository. Add a collect hook that reads each scenario's spans from the tracing backend and returns the tool calls it finds in the calls field, with the trace URL as an extra field. Ask me for the backend's query details instead of guessing, and show the proposed hooks and the environment-variable names they need. Before you create or test anything, list the writes a test call could make in staging and wait for my approval.
Note

Note: Test how your agents actually behave across workflows, tools, and actions. Get started with Agent Assurance

Conclusion

Start your AI agent tracing review with propagation: search your tracing backend for tools/call server spans that sit at the root of their own trace. Each one is a hop that dropped context, or a caller you do not trace.

Then pick one write your agent makes and check it against the system it writes to. TestMu AI's Agent Assurance quickstart uses a public sample agent, so you can run one small test before you connect an agent in your own staging.

Author

...

Sandeep Yadav

Blogs: 10

  • Linkedin

Sandeep Yadav is a Senior Software Engineer at TestMu AI (formerly LambdaTest), where he builds the platform's test intelligence and AI-native engineering systems. He has architected autonomous GitHub Apps, vector-search code intelligence, and self-diagnosing QA workflows, and designed distributed platforms that process 2M+ daily test executions and 1B+ events, turning high-volume test, log, and code data into intelligent, self-optimizing systems. He works on embedding reasoning models into production infrastructure to power autonomous review, root-cause analysis, and analytics workflows. He brings over four years of engineering experience with deep expertise in the Elastic Stack, Apache Kafka, and Redis. Earlier he engineered a GDPR-compliant, end-to-end-encrypted secure web-chat application at Mithi. A Facebook Hackercup 2021 Round 2 qualifier and merit-scholarship recipient, Sandeep holds a B.Tech in Electrical Engineering from Delhi Technological University.

Reviewer

...

Samyak Goyal

Reviewer

  • Linkedin

Samyak Goyal is a Senior Member of Technical Staff at TestMu AI engineering Kane CLI, the command-line tool that runs browser automation from the terminal, where a flow described in natural language executes in a real Chrome browser and returns pass or fail with shareable proof. He is a backend engineer with 4+ years of experience, previously an SDE at Innovaccer, where he built APIs, introduced Kafka, and cut deployment from weeks to hours. Samyak also builds multi-agent systems, skill-orchestration frameworks, and a personal copilot that indexes 200+ microservice repositories.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

AI Agent Tracing FAQs

Did you find this page helpful?

More Related Learning Hubs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests