The situation
The subscriber help desk runs on an agent hosted in the AgentCore runtime, reaching its tools through a gateway. A message comes in, the agent reasons about it, calls a tool to fetch the account, checks a knowledge base for the refund policy, calls another tool to compute a figure, and writes a reply. Most of the time it is right. This morning it told a subscriber they were owed nothing when they were owed a fortnight’s box, and the only thing anyone can see is that closing sentence.
The final answer is the one artefact that carries none of the information you need. A wrong number at the end could come from a bad tool argument, a stale knowledge-base document, a correct tool result the model then used wrongly, or a step it never took. From the reply alone you cannot tell which. So you cannot fix it, and you cannot tell the subscriber what went wrong with any honesty.
An agent run is a chain of decisions: read the request, pick a tool, pass it arguments, read what comes back, move to the next step, repeat until done. Debugging one, or proving to an auditor what it did, means reconstructing that chain for a single request. What matters is what to capture, where it lands, and how you pull one run back out of a day’s traffic.
What actually matters
Tracing is a different job from evaluation and from guardrails, and the three are easy to conflate. Tracing answers “why did it do this”, reconstructing one run step by step. Evaluation answers “is it any good”, scoring outputs across many runs against references or a judge model. Guardrails block or filter content before it reaches anyone. A guardrail that blocked a response tells you nothing about the steps that led there, and an evaluation score tells you the agent is worse this week without naming the step that changed. Only a trace reconstructs the decision path.
The second thing is the split between the model’s reasoning and the plumbing around it. The reasoning layer is the run the agent actually took: each turn, the tool or knowledge base it called, the arguments it sent, and what came back. That is the layer that explains why. Underneath, each tool is usually a Lambda calling real systems, and that layer explains what happened when the tool ran: the API it hit, the latency, the error it turned into a default. A bad tool call and a bad step of reasoning look identical from the final answer, and completely different once both layers are captured.
The third is verbatim capture versus summary. A trace gives you the shape of the run. For a real audit or a subtle bug you often need the exact prompt the model was sent and the exact completion it returned, byte for byte. Model invocation logging records those. A trace tells you the model called the billing tool; the invocation log carries the precise text it generated to do so.
The fourth is production reach: getting to one run out of thousands, being alerted when something drifts, and keeping the record long enough to matter. A trace you can only see by re-invoking the agent with tracing switched on is fine for a repro and useless for the incident that already happened. What you want is traces, metrics and logs landing in a store you can query afterwards, with retention you control, tied together by an identifier. All of that is configured before the run and never after, and the run you most want to read is always one that happened before you got around to it.
What we’ll filter on
- Can you reconstruct a single run end to end: the reasoning, each tool call, its arguments, and what it returned?
- Are the exact prompts and completions captured verbatim, not just summarised?
- Does the picture join up across the agent and the Lambda tools underneath it, or does it stop at the agent boundary?
- Is it queryable and alertable in production after the fact, or only visible while you re-run the agent?
- How long is it retained, and can it stand up as an audit record of what happened?
The landscape
These are layers that stack, not rivals you choose between. A production setup usually runs several at once, and the first thing to know is how little of it is on by default.
AgentCore’s built-in metrics
Out of the box, AgentCore publishes metrics to CloudWatch for every resource type: the runtime, memory, the gateway, the built-in tools, and identity. Session count, latency, duration, token usage and error rates all arrive without you writing anything, and they show up on the CloudWatch generative AI observability page.
This is the layer that tells you something is wrong. It is aggregate by nature, so it will show the error rate climbing at ten past nine and say nothing about which subscriber, which tool, or which argument. Useful for alarms, useless for the single run you have been asked to explain.
Spans and traces from an instrumented agent
This is the layer that explains why, and it has to be switched on deliberately. AgentCore’s model has three tiers. A session is the whole interaction with one subscriber. A trace is a single request-response cycle inside it. A span is one unit of work inside that, with a start, an end, a status and a parent. A run therefore comes out as a tree: the invocation at the top, reasoning turns, tool calls and knowledge-base lookups beneath it. Walking that tree back to the step where the logic went wrong is what separates a misapplied tool result from a wrong step of reasoning.
Getting it needs two things that are easy to miss. CloudWatch Transaction Search has to be enabled once for the account, and until it is, you cannot enable tracing on an AgentCore resource at all. Then the agent has to be instrumented with the AWS Distro for OpenTelemetry: add aws-opentelemetry-distro to the dependencies and run the agent through opentelemetry-instrument. Service-provided spans are a per-resource toggle on top of that. Only the metrics arrive without one.
Where those spans end up is AWS X-Ray, with Transaction Search as the search surface over them. An agent instrumented with ADOT and emitting OpenTelemetry spans is how one run becomes a trace you can pull up by id a week later. Ordinary request logging carries none of the reasoning, which is why agent work needs its own pipeline: spans shaped by the GenAI semantic conventions, with each prompt version tied by trace id to the runs it produced.
If the agent is built on Strands, LangChain or CrewAI, the framework already speaks OpenTelemetry and the GenAI semantic conventions, so auto-instrumentation carries most of the load. Spans land in the agent’s own log group, /aws/bedrock-agentcore/runtimes/<agent_id>-<endpoint_name>, which keeps spans, logs and standard output together per agent. That is the default for newly created agents, and it needs ADOT 0.18.0 or later; older agents and older distro versions deliver to the shared aws/spans group instead.
Model invocation logging
A Bedrock setting, enabled per account and Region and unrelated to AgentCore, that captures the input and output of every InvokeModel and Converse call and delivers them to CloudWatch Logs, S3, or both. This is the verbatim record: what the model was sent and what it returned, byte for byte, up to 100 KB a body. Larger bodies and binary payloads go to S3 as separate objects instead.
It is the backbone of an audit trail, and what you reach for when a span’s attributes are not precise enough to explain a subtle failure. Because it stores completions byte for byte, it also supports output diffing: pull the responses to the same prompt across a fortnight and compare them, and a template edit that changed behaviour shows up as a difference in the recorded text. It records whatever was in the prompt, personal data included, so retention, encryption and access controls are part of turning it on rather than a later tidy-up.
Reading the record back with Logs Insights
Once the invocation log is landing in CloudWatch Logs, CloudWatch Logs Insights is where you analyse prompts and responses. Two query shapes cover most of an incident. The first works over the application’s own log group, where the wrapper around the model call records a model id, a status and a latency:
fields @timestamp, modelId, errorCode, latencyMs
| filter modelId = 'anthropic.claude-sonnet-4-5-20250929-v1:0'
| stats count(*) as calls, pct(latencyMs, 95) as p95 by errorCode, bin(5m)
The second goes to the invocation log group and lifts one request out of it:
fields @timestamp, input.inputBodyJson, output.outputBodyJson
| filter requestId = 'c3f0a1e2-8c2f-4d19-9f77-2b6a55e0d431'
| limit 1
The first tells you which model calls are slow or failing, and in which five-minute window. The second shows the prompt that model was sent and the completion it returned. Exact exception names matter in that first query, because the common failures look alike in the console and have different fixes: a ValidationException from a malformed tool schema or from an input past the model’s context limit, a ThrottlingException from a burst of concurrent sessions, a ModelTimeoutException from a call that ran past the model timeout.
Distributed tracing into the tool Lambdas
The gateway targets behind this agent are Lambda functions calling downstream systems, and their execution is invisible from the agent’s own spans. Add the AWS Lambda Layer for OpenTelemetry to each one and set AWS_LAMBDA_EXEC_WRAPPER to /opt/otel-instrument, and the function auto-instruments: the services it touched, the latency of each hop, and where an exception was thrown. The layer reports into AWS X-Ray, the same service the agent’s spans land in, and Transaction Search ingests both, so a tool’s subsegments sit under the agent’s spans in one trace. Auto-instrumentation also wraps the AWS SDK client, so a Bedrock InvokeModel or Converse call made inside a tool arrives as its own subsegment. A slow model call then sits next to a slow DynamoDB query in the same waterfall, which is the property you want when the agent, the gateway and three Lambdas are separate services with separate logs.
One gotcha worth carrying: the ADOT Collector is not supported for agent observability. It is the ADOT SDK or the Lambda layer, and reaching for the collector out of habit produces telemetry that never arrives.
Correlation, which is what makes any of it usable
Two identifiers do the stitching. The session id travels on the X-Amzn-Bedrock-AgentCore-Runtime-Session-Id header and groups a subscriber’s whole conversation. The trace id travels as X-Amzn-Trace-Id in X-Ray format, or as the W3C traceparent header, and groups one request across the agent and everything it called.
Propagate both from the front end inward and “show me what happened to subscriber 8c2f at 09:10” is a query. Skip them and the same data is scattered across log groups with no way to know which rows belong together, which is most of the difference between an observability bill and an observability capability.
From spans to tool metrics
The spans that debug one run become a tool performance picture as soon as you aggregate them. Every tool call already carries a name, a duration and a status, so rolling them up gives per-tool call volume, p50 and p99 duration, error rate, retry rate, and how often the model calls each tool relative to the others. Emit those as CloudWatch metrics from the span attributes, or straight from the tool Lambdas with embedded metric format, and they sit on the same dashboards as the rest of the application rather than in a query someone writes by hand mid-incident.
That turns “the agent feels slower this week” into “calcRefund is being called four times a session instead of once”. Set a baseline for each tool and put a CloudWatch anomaly-detection band on its call count, and you catch the failure that trips no alarm anywhere else: a prompt edit that made the model over-call a tool, running at four times the volume until it shows up on the bill.
Coordination between agents is the same instinct one level up. Where a supervisor hands work to sub-agents, propagate the session id across the supervisor and every sub-agent so a trace shows the whole handoff chain. Then track handoff count, handoff depth, and loop detection, the same pair passing control back and forth more than twice in one session, as metrics in their own right. A coordination failure has no slow span in it: every hop looks normal, and the session takes nine seconds because it made eleven hops.
Not tracing, but next to it
AgentCore Evaluations scores agent quality from the sessions, traces and spans you are already collecting. Bedrock evaluations score models and knowledge bases across many runs, and Bedrock Guardrails block or filter content in flight. All three produce useful signals and none of them reconstructs a decision path. Keep them in the mental map so you do not start an evaluation job when what you need is one run’s spans.
Evaluation
Side by side
| Signal | Reconstructs one run | Verbatim prompts and completions | Reaches the tool Lambdas | On without setup | Audit-grade retention |
|---|---|---|---|---|---|
| Built-in AgentCore metrics | ✗ (aggregate) | ✗ | ✗ | ✓ | ✓ (as configured) |
| Instrumented spans and traces | ✓ | ✗ (attributes, not full text) | ✓ (with instrumented tools) | ✗ (Transaction Search + ADOT) | ✓ (as configured) |
| Model invocation logging | Partial (per model call) | ✓ (to 100 KB inline) | ✗ | ✗ (account and Region setting) | ✓ |
| Distributed tracing (Lambda layer) | Partial (the tool side) | ✗ | ✓ | ✗ (per function) | ✓ (as configured) |
| Aggregated tool metrics | ✗ (aggregate) | ✗ | ✓ | ✗ (from spans or EMF) | ✓ (as configured) |
| Evaluation / Guardrails | ✗ | ✗ | ✗ | ✗ | n/a |
Reading it for this incident: the spans tell you which steps the model took, invocation logging gives you the exact text it generated, and tool tracing tells you whether the billing Lambda returned what the model acted on. The debuggable, auditable setup is spans for the reasoning, the invocation log for the verbatim record, and tool tracing for the plumbing, correlated by session and trace ids. Notice which column is nearly empty: only the aggregate metrics arrive without setup, and they are the one signal that cannot explain a single run.
An agent run as a trace
The solution
Spans, for the reasoning. Enable CloudWatch Transaction Search once for the account, add the ADOT SDK to the agent, and every run thereafter is stored as a span tree: a parent for the invocation, children for each reasoning turn, tool call and knowledge-base lookup, with timing and status on each. This is the signal you read first when an answer is wrong for no obvious reason. Its limit shapes how you use it. It describes the agent’s view, so it shows that the model called calcRefund with a particular weeks argument and says nothing about what happened inside the Lambda that served the call. Because the spans are stored rather than attached to a live invocation, the run that already failed is still there to read.
Model invocation logging, for the verbatim record. Enable it once per account and Region, and every model call thereafter delivers its request and response to CloudWatch Logs or S3. When a trace summary says the model returned no refund and you need to know exactly what it generated to get there, the invocation log has the literal text. It is precise, it is complete, and it lives in storage you control with retention you set. Treat it accordingly: it records prompts and completions verbatim, which can include personal data, so it needs tight access controls, sensible retention and encryption. Plan for its volume rather than discover it on the bill.
Built-in metrics, for health and alarms. AgentCore’s default metrics (session count, latency, duration, tokens, errors) and your Lambda logs are where you watch the fleet and get told when something moves. You alarm on an error-rate spike or a latency climb and it points you at a window, then you pull the individual traces and invocation logs for that window. Metrics start investigations; they do not finish them.
Tool tracing, for the plumbing. Add the OpenTelemetry Lambda layer to the functions behind the gateway targets and each tool execution becomes a trace of its own: the downstream services it called, the latency of each, and where an exception was thrown. Correlate those with the agent’s spans and the two failure modes separate. “The model passed the wrong argument” shows up in the agent’s spans; “the tool returned a default because the billing service timed out” shows up in the tool’s. Propagating the session and trace ids from the invocation into the tool calls is what lets you stitch one request together across both.
Tool metrics, for the pattern no single run shows. Aggregate the tool spans into per-tool volume, duration percentiles, error rate and selection distribution, with an anomaly band on each tool’s call count, and you get the layer that says which tool and which window. The spans then say why. A tool whose p99 doubled on Tuesday, and a tool being called four times a session, are both invisible in any one trace and obvious in a fortnight of them.
The identifiers, because they are what makes it a system. Send the session id on X-Amzn-Bedrock-AgentCore-Runtime-Session-Id and a trace id as traceparent from the front end, and propagate both inward. That turns “show me run 8c2f from this morning” into a query across stored spans rather than a hunt through log groups, and it lets a span in the agent and a segment in a tool Lambda be recognised as the same request. Because the telemetry is OpenTelemetry-shaped, the same identifiers work if you send it somewhere other than CloudWatch, which is a matter of setting DISABLE_ADOT_OBSERVABILITY=true and pointing the exporter elsewhere.
Worked example
The subscriber wrote in; the agent replied that no refund was due; the subscriber was owed two weeks. The reply is the only thing the help desk could see, so the investigation starts by pulling the run.
The span tree lays out the chain. Turn one: fetch the subscription, getSubscription(id=8c2f), returning paused=true. Turn two: consult the refund policy knowledge base, two chunks returned about the refund window. Turn three: compute the refund, calcRefund(id=8c2f, weeks=?). Turn four: compose the reply. The tree shows the shape and already narrows the field: the model did reach the calculation step, so no step was skipped. The suspect is the weeks argument it passed.
The invocation log for that model call has the verbatim completion, and there it is. The model generated weeks=0, and the text it produced counts only the current week, when the subscriber had been paused for two delivery cycles. The tool then behaved exactly as specified: a zero-week refund computes to zero.
To confirm that, the tool trace for calcRefund, correlated by trace id, shows the Lambda receiving weeks=0, calling the billing service, and getting AUD$0.00 back in 40ms with no error and no timeout. The plumbing was healthy. The defect was upstream, in how the pause history became an argument.
That is a fix you can make with confidence: the pause data the model is given is ambiguous about multi-cycle pauses, so you tighten what the account tool returns and add an instruction about counting cycles. Without the trace you would have been guessing between a tool bug, a stale policy document and a reasoning error. With it, the wrong step is named, the verbatim proof is on file, and the tool is ruled out. If an auditor asks later what the agent did for subscriber 8c2f on this date, the same three artefacts answer them.
What’s worth remembering
- The final answer is the one artefact that hides the reason; debugging or auditing an agent means reconstructing the run behind it, not reading the reply.
- Stored spans are the layer that explains the model’s reasoning, and on AgentCore they take setup: Transaction Search for the account, ADOT in the agent, and tracing enabled per resource.
- Model invocation logging is the verbatim record, delivered per account and Region to your own logs or S3, with bodies over 100 KB landing in S3 as separate objects.
- The agent’s spans stop at the agent; tracing the tool Lambdas shows what they did, and correlating the two by session and trace id separates a bad argument from a bad tool.
- Only the aggregate metrics arrive without setup, and they are the one signal that cannot explain a single run; everything that can is configured before the run, and none of it is retroactive.