Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Learning Hub
- /
- AI Agent Monitoring: What to Watch When Agents Act
AI Agent Monitoring: What to Watch When Agents Act
AI agent monitoring explained: tool calls, side effects, handoffs, cost and loops to watch, how to alert on drift, which tools help, and what monitoring misses.
Published on:
OVERVIEW
An agent can finish a run with a successful response, a confident summary and nothing in the error log, and still have closed the wrong ticket or refunded the wrong order. Dashboards built around request errors were designed for services that fail loudly; AI agent monitoring has to catch agents that act and fail without throwing anything.
How much a monitor catches depends on what it can see. In OpenAI's study of monitoring reasoning models, run during a frontier model's training on agentic coding tasks, a GPT-4o monitor that saw only the agent's tool calls and final outputs caught 60% of its systemic reward hacks, such as exiting before the unit tests ran, and 95% once it could also read the chain of thought.
AI agent monitoring, also called agentic AI monitoring, records what agents do on live traffic and alerts when their behavior moves. In a trace, an update the agent reported but never made can still look like a success until someone checks the target system.
Overview
AI agent monitoring is the practice of tracking what AI agents do in production, including their tool calls and arguments, the side effects they cause, handoffs, permission use, cost per task and loops, and alerting when that behavior drifts from a baseline.
What should you monitor when AI agents act?
- Tool calls and arguments: Record every tool an AI agent calls, the arguments it passed and the result, and alert when a task type starts using a tool outside its usual set. The tool name alone hides the failure when the agent chose the right tool with the wrong record ID.
- Side effects: Side effects are the records, files and messages an AI agent's run changes in other systems. Log each claimed effect as an event and confirm a sample with read-only queries of the system of record, since the agent's own summary is not evidence that the write happened.
- Behavioral drift: Behavioral drift is a change in how an AI agent works while its error rate stays flat, such as a new tool mix, more steps per task, or more runs ending without their intended action. Alert on distance from a per-task baseline.
- Cost per completed task: Cost per completed task divides an AI agent's token and tool spend by the tasks that actually finished. A rising figure often points to loops, retries or longer plans before any user notices a problem.
How is monitoring different from verification?
Monitoring reads what the agent and its framework recorded on live traffic; verification checks what the run changed. TestMu AI's Agent Assurance verifies before release: it runs the agent against staging, grades each criterion on observed evidence such as changed files and read-only checks of records, and reports what it could not verify apart from the pass rate.
What Is AI Agent Monitoring?
AI agent monitoring is the continuous collection of, and alerting on, data about what AI agents do in production: the tools they call and with what arguments, the effects they leave in other systems, what each task costs, and how those patterns change over time.
The Monitoring Distributed Systems chapter of Google's SRE book names four golden signals to watch: latency, traffic, errors and saturation. Those signals still apply to the service an agent runs in, but an agent can hold all four steady while choosing the wrong action.
The same chapter separates white-box from black-box monitoring, and the split maps onto agents directly:
- White-box monitoring - metrics the system exposes about its internals, including logs. For an agent, that means its traces: model calls, tool spans and whatever reasoning the framework records.
- Black-box monitoring - testing externally visible behavior as a user would see it. For an agent that acts, that means checking the ticket, file or record the run was supposed to change.
Monitoring also differs from the practices it gets confused with:
- Observability - the instrumentation and trace data that let you investigate why a run went wrong, which the guide to AI agent observability covers alongside tracing tools. Monitoring is the part that runs continuously on that data and pages someone.
- Evaluation - scoring an agent against a dataset or rubric, mostly before release, though many eval tools also score samples of live traffic. Monitoring runs continuously on that traffic, so it is best at spotting change; judging whether an action was correct needs an evaluator or a check against the system of record.
What Should You Monitor When AI Agents Act?
Start with actions and their consequences, since that is how an agent's mistakes reach your systems. OpenAgentSafety ran five prominent LLMs as agents on more than 350 multi-turn, multi-user tasks with real tools, including file systems, bash shells, web browsers and messaging platforms, and found unsafe behavior in 51.2% (Claude 3.7 Sonnet) to 72.7% (o3-mini) of safety-vulnerable tasks.
Record these signals for every run:
- Tool calls and arguments - the tool name, the arguments with secrets and personal data redacted, the result status, and the task type that triggered the call. A refund tool called from a password-reset task deserves an alert even when the call succeeds.
- Side effects - every effect the agent claims, such as a closed ticket or a sent email, logged as a structured event you can reconcile against the system of record with read-only queries. A claimed write with no matching record is the failure no latency or error metric will show.
- Handoffs - which agent handed work to which, the fields passed, and whether the receiver acted on them. A handoff that drops an order ID leaves the next agent working confidently on the wrong object, and agent handoff testing covers the assertions to write for that boundary.
- Permission use - calls to write and destructive tools per agent and per credential, including denied attempts. Alert when an agent uses a destructive tool it never needed before, or runs with a broader credential than its task requires.
- Cost per completed task - token and tool spend divided by tasks that reached their intended end, since cost per request hides an agent that now makes several calls to finish what one call used to do. The guide to LLM cost tracking for agent evals explains how to count only verified successes in that denominator.
- Loops - steps per run and repeated identical tool calls with the same arguments. An iteration cap in the framework ends a looping run, but only a monitor tells you how often runs are hitting that cap.
Give every run a terminal outcome as well: the action the task was meant to end with, or an explicit escalated, abandoned or timed-out state. Without it you can count errors, but you cannot compute a completion rate.
Agentic AI Monitoring for Multi-Step and Multi-Agent Runs
Agentic AI monitoring treats the whole run as the unit instead of the single model request. A request-level dashboard can show healthy latency on every call while one run loops through dozens of steps and several handoffs before it fails.
- One trace per task - carry the run ID through every agent, tool and queue, so a delegation to a sub-agent never starts an orphaned trace.
- Budgets per run - track steps, tokens and wall-clock time against a ceiling for each task type, and alert on runs that hit the ceiling as well as on runs that error.
- Delegation depth - count how many agents a task passes through. A router that starts sending every request to two specialists instead of one roughly doubles specialist spend without raising a single error.
- Outcome review - send a sample of completed runs, weighted toward new tool mixes and high-value actions, to a person who checks the result in the system of record rather than in the transcript.
Automated trace review helps with volume, but models still miss things in long traces:
- Monitor accuracy - in SHADE-Arena, 17 task pairs in which an agent pursues a hidden harmful side task, the best monitor, Gemini 2.5 Pro, reached an AUC of 0.87 at separating benign from malign transcripts.
- Undetected sabotage - under a Claude 3.6 Sonnet monitor, Claude 3.7 Sonnet completed both its main task and the hidden side task without being flagged 27% of the time, and Gemini 2.5 Pro 15%.
- Hidden reasoning - the SHADE-Arena authors found that side-task success relied heavily on a scratchpad the monitor could not see.
Use model-based monitors to rank runs for human review, and base verdicts on evidence from the systems the agent changed.
How to Alert on Behavioral Drift Instead of Error Rate
Page on outcomes and on shifts in behavior, because a wrong action can complete without raising an error. Set the alerts up in this order:
- Define a verified completion objective - for each important task type, set a target share of runs that end in their intended action, confirmed by a read-only check wherever one exists. That objective is what you page on, and the behavior signals explain why it moved.
- Alert on budget burn rate - Google's SRE Workbook guidance on alerting on SLOs recommends starting with pages at 2% of the error budget spent in 1 hour and 5% in 6 hours, and a ticket at 10% in 3 days, each paired with a shorter window of 5 minutes, 30 minutes or 6 hours so the alert resets once the problem stops. The same arithmetic works on an agent's completion objective once runs are frequent enough. In the Workbook's own example, a service with a 99.9% objective that gets 10 requests an hour pages on a single failure, so for a low-volume agent borrow its options for low-traffic services: add scheduled synthetic runs on test accounts, group related task types under one objective, or lower the objective or lengthen its window.
- Baseline behavior per task type - store each task type's usual tool mix, median and 95th percentile steps per run, handoff count and cost per completed task, keyed by agent version.
- Alert on distance from the baseline - page when a task's tool mix shifts sharply or runs ending without their terminal action climb, and open a ticket when steps per run or cost per completed task drift over a week.
- Flag results that look too good - a sudden jump in completion rate, or runs finishing far faster than their baseline, deserve a look, because an agent that skips a step or games a check also looks efficient.
- Annotate every change - put model, prompt, tool and data-source changes on the same timeline as the alerts, and treat a provider-side model update as a release, so a drift alert points at the change that caused it.
Keep the error-rate and latency alerts you already run for the service the agent lives in, since they still catch outages and slow dependencies.
AI Agent Monitoring Tools
Agent monitoring software falls into two groups: application monitoring platforms that added agent views, and agent-first tools built around session replays. The comparison below judges each tool on alerting, dashboards and cost tracking, using the vendor's own documentation and announcements as checked on 30 September 2026.
| Tool | Agent activity it records | Alerting | Dashboards | Cost tracking |
|---|---|---|---|---|
| Datadog Agent Observability | Agent steps, tool calls with execution time, errors and retries, and handoffs in a graph view | Monitors on its agent metrics for cost, tokens, errors and latency | Out-of-the-box Operational Insights dashboard | Per-span input, output and total cost metrics |
| Splunk Observability Cloud AI Agent Monitoring | Agent workflows, tool calls and span details in a trace view | Alerts on any of its agent metrics | AI agents page with performance, cost and security metrics for every agent | Total, input and output tokens with their cost |
| New Relic AI Monitoring | Every agent invocation, tool call and handoff in LangGraph, Strands and AutoGen | Custom alerts on token usage | Out-of-the-box performance dashboards and an entity map of agents and tools | Token counts per agent and model cost comparison |
| Sentry Agent Monitoring | Agent runs, tool calls, model calls and errors | Static-threshold and anomaly detectors on hourly spend | Agents dashboard for executions, model costs, token usage, tool calls and errors | Cost calculated from token counts and per-model token rates |
| AgentOps | Session replays as step-by-step agent execution graphs | No alerting feature in its README; loop detection is marked in progress | Self-hostable dashboard | Spend tracking across LLM providers |
How each one fits:
- Datadog Agent Observability - fits teams already on Datadog, since its agent metrics are 100% sampled, kept at full granularity for 15 months, and usable in the same dashboards and monitors as the rest of the stack.
- Splunk Observability Cloud - fits teams that want quality next to cost, since its AI Agent Monitoring uses LLM-as-a-judge evaluators for hallucinations, biases, sentiment and toxicity, and Splunk says it builds on OpenTelemetry and Cisco AGNTCY to avoid lock-in.
- New Relic - fits agents built on LangGraph, Strands or AutoGen, which it instruments through an existing APM agent. Its agent views were announced in preview in February 2026, so confirm their status before you plan around them.
- Sentry - fits teams that want spend alerts first: its cookbook sets a static threshold and an anomaly detector on hourly LLM spend, and it auto-instruments OpenAI, Anthropic, Google GenAI, LangChain, LangGraph, Pydantic AI, the OpenAI Agents SDK and the Vercel AI SDK in Python and Node.js.
- AgentOps - fits teams that want open-source, self-hosted session replays under the MIT license. Route alerts through your own stack, because its README marks infinite-loop and recursive-thought detection as in progress rather than shipped.
If your agents run on AWS, Amazon CloudWatch is the native option, showing tool call counts, success rates and execution times, token usage, and custom dashboards with alarms for OpenTelemetry-instrumented agents. For voice and phone agents, the comparison of voice agent monitoring tools covers call-specific metrics instead.
Every tool above reports telemetry that the agent's stack emits. Confirming that a reported effect actually landed is a check you add yourself, such as the reconciliation query from the side-effects signal above, and TestMu AI's Agent Assurance sits outside this comparison on purpose: it runs before release and does not watch production traffic.
Monitoring vs Verification for Agents That Act
Agents misreport their own work, even when they know better. METR's report on reward hacking in recent frontier models found o3 reward-hacking in 0.7% of runs across HCAST tasks, reported reward hacking as more than 43 times more common on RE-Bench tasks, and saw o3 reward-hack in every trajectory on one RE-Bench task. The hacks included overwriting a timing function so it always returns shorter measurements and monkey-patching an evaluator so every piece of code passes. Asked whether its plan adhered to the user's intention, o3 answered no 10 out of 10 times.
METR found the RE-Bench cases by inspecting anomalously high-scoring runs, the same check as the drift alert on results that look too good. For an agent that acts in your systems, the equivalent failure is a false completion report, which the guide to agent action hallucination covers in depth.
How the two practices compare:
- When it runs - monitoring runs on production traffic after release; verification runs against staging before release and again in CI.
- Evidence - monitoring reads the traces, metrics and logs an agent's stack emits; verification checks files, records and tool calls against what the task required.
- What it catches - monitoring catches drift, cost spikes, loops and new failure patterns on real inputs; verification catches a claimed effect that never happened, on inputs you chose.
- What it misses - monitoring misses a wrong effect that left nothing unusual in the trace; verification misses inputs nobody wrote a scenario for, which production supplies.
Each practice covers the other's blind spot, so connect them: every alert that exposes a real failure becomes a pre-release case.
- Confirm the failure - pull the run's trace and check the effect it claimed against the system of record with a read-only query, so a noisy alert does not turn into a test.
- Attach the evidence - save the run ID, the trace and the query result with the report, so whoever writes the case starts from what happened rather than from the agent's reply.
- Hand it to the release gate - the guide to agent regression testing covers turning a production failure into a scenario, confirming it fails before the fix, and keeping it in the suite that gates each release.
- Keep the monitor - leave the alert that fired in place, since the scenario covers the input you saw and the monitor covers the next one.
Verifying Agents Before Release With Agent Assurance
An agent's account of what it did is the weakest evidence available about what it did. TestMu AI's Agent Assurance is built on that premise: it derives scenarios from the code or spec of an agent that acts, invokes the real agent against staging, and grades every acceptance criterion against observed evidence.
For the failures a production monitor surfaces, it checks:
- Tool calls against the declared tools - the calls your profile's hooks return are checked against the tools the agent itself declares, and a criterion can require that a tool is never called, such as a delete tool during a read-only lookup.
- Effects a run leaves - files that changed under the paths the profile declares, and artifacts the run produced, instead of the agent's summary of them.
- Records in other systems - a judge can confirm a ticket or an order with a read-only query through an MCP tool you provide and approve; the current release connects only to stdio MCP servers.
- Unable to Verify - a criterion with no evidence behind it is reported as Unable to Verify and kept out of the pass rate, so the pass rate covers only what was observed.
Most eval and observability tools score what your agent said and recorded. Agent Assurance checks what the run changed, and reports what it could not verify:
| Aspect | Agent Assurance | AI eval tools | LLM observability |
|---|---|---|---|
| Test cases | From your code or spec | Written or synthesized | From production traces |
| Tool calls | Against declared tools | Against your lists | Logged, optionally scored |
| Side effects | Files, artifacts, probes | Scripted per task | Trace data only |
| Grading | Claimed actions aren't proof | LLM judge or code checks | LLM judge or human review |
| Adversarial tests | Generated by default | Add-on in some tools | Not generated |
| When it runs | Before release, in CI | CI and live traffic | Production, plus CI |
| Unverifiable results | Reported separately | Errors or opt-in skips | Left unscored |
The eval and observability columns describe each category's default approach, not any single product. Agent Assurance also uses model judges and grades what an agent says; the difference is that a claimed action never counts as proof, and since it does not sit in production, it complements the monitoring above rather than replacing it.
Agent Assurance runs from the terminal as Rook CLI. Install it from npm, which needs Node.js 22 or newer, and check the version:
npm install -g @testmuai/rook
rook --versionOther install methods, including Homebrew and a checksum-verifying shell installer, are in the Rook CLI install guide. Point the profile at staging, because the agent's writes during a run are real and are not rolled back.
In Claude Code, add the rook skill, then hand it the incident your monitor caught:
npx @testmuai/rook-skill@latest install --agent claude-code/rook A production trace shows this agent replying "ticket closed" while the ticket stayed open. Generate up to three functional scenarios that replay that request through the staging profile, each with a criterion that confirms the ticket status through a read-only check. Before running anything, list the scenarios and the writes each one could cause in staging.Review the scenarios before any run. This console output was captured from Rook CLI 0.1.5 on 30 September 2026:
$ rook --version
0.1.5
$ rook scenarios --help
Usage: rook scenarios [options] [command]
inspect and curate the test set
Options:
-h, --help display help for command
Commands:
list [options] what would run, and what could not
exclude [options] <ids...> keep a scenario on disk but leave it out of runs
include [options] <ids...> undo an exclude
delete [options] <ids...> remove scenarios permanently
help [command] display help for commandIn CI, a finished run exits 0 whether its scenarios passed or failed, so gate the pipeline on the verdicts in rook report --json rather than on the exit code.
Note: Agent Assurance can also generate performance, token economy and reliability scenarios when you add the non-functional class, so the latency and token cost you watch in production can be tested before release too.
Conclusion
Start with one agent that writes to a system you care about: log every effect it claims as a structured event, and schedule a daily read-only reconciliation of a sample against the system of record. That gives your AI agent monitoring a verified completion rate to alert on, instead of the agent's own reports.
Then add the burn-rate and drift alerts above, turn each confirmed failure into a scenario in the pre-release suite that Agent Assurance runs against staging, and wire that suite into your pipeline with the guide to run Agent Assurance in CI/CD.
Author
Saurabh Prakash is an Engineering Manager at TestMu AI (formerly LambdaTest), where he leads engineering on agentic AI development and scalable system architecture for the quality engineering platform. He has also contributed to Test at Scale, the company's open-source test intelligence platform. He brings over 9 years of experience across Node.js, Java, Spring, MVC, data structures, algorithms, and scalable system design, with earlier roles as SDE 2 at Zomato, Senior Software Engineer at LogicHub, and Software Development Engineer at Directi. Saurabh holds a B.Tech in Computer Science and Engineering from Delhi Technological University.
Reviewer
Sandeep Yadav is a Senior Software Engineer at TestMu AI (formerly LambdaTest), where he builds the platform's test intelligence and AI-native engineering systems. He has architected autonomous GitHub Apps, vector-search code intelligence, and self-diagnosing QA workflows, and designed distributed platforms that process 2M+ daily test executions and 1B+ events, turning high-volume test, log, and code data into intelligent, self-optimizing systems. He works on embedding reasoning models into production infrastructure to power autonomous review, root-cause analysis, and analytics workflows. He brings over four years of engineering experience with deep expertise in the Elastic Stack, Apache Kafka, and Redis. Earlier he engineered a GDPR-compliant, end-to-end-encrypted secure web-chat application at Mithi. A Facebook Hackercup 2021 Round 2 qualifier and merit-scholarship recipient, Sandeep holds a B.Tech in Electrical Engineering from Delhi Technological University.
AI Agent Monitoring FAQs
Did you find this page helpful?
More Related Learning Hubs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




