Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Pre-Production vs Post-Production: Where AI Agent Testing Belongs Now That Observability Does Evals
Pre-Production vs Post-Production: Where AI Agent Testing Belongs Now That Observability Does Evals
Observability vendors now score AI agents in production. Agent Assurance tests them before release. Here is where each fits and how they work together.
Published on:
AI agents need two kinds of checks. Pre-production agent testing determines whether an agent is fit to ship by running it against scenarios and grading the changes it actually makes. Post-production observability scores how the agent behaves on live traffic.
Agent Assurance from TestMu AI (Formerly LambdaTest) owns the first; Dynatrace, Splunk, and New Relic are moving into the second.
Observability Vendors Just Moved Into Evals
Between April and October 2026, three of the largest observability vendors moved to bring agent evaluation in-house:
- Dynatrace completed its acquisition of Arize, announced at $915 million, on October 1, 2026, and says Arize's LLM and agent evaluation will be integrated into its observability platform over time.
- Cisco acquired Galileo, which has been available since September 2026 as Splunk Agent Observability, with out-of-the-box evaluators and runtime guardrails.
- New Relic announced AI Evaluation on October 6, 2026. It uses an LLM judge to score AI responses from development to production, and opens in public preview in November.
This is good news for anyone shipping agents. It confirms that a dashboard of latency and error rates cannot tell you whether an agent applied the right refund policy. Every request can return a 200 while the agent does the wrong thing.
It also raises a fair question for engineering teams: if my observability tool now runs evals, do I still need agent testing? Yes, because the two answer different questions at different times.
What Post-Production Observability Does Well
Observability evals what already happened. They read production traces (model calls, retrievals, tool calls, handoffs) and attach quality scores, often with an LLM judge. That makes them the right tool for:
- Drift: Is the agent's reasoning changing week over week on real traffic?
- Cost and latency per run, tied back to the services and infrastructure involved.
- Runtime guardrails that block an unsafe output as it happens.
- Root cause when something breaks in front of a customer.
Several platforms also run LLM evals in CI: Arize and Galileo both run experiments against a dataset you build, and the roundup of AI evaluation tools for production compares more of them.
The guide to AI agent monitoring covers what to alert on, and the comparison of AI agent observability tools shows what each platform records for a multi-step run.
Observability is the system of record for an agent in production. Keep it.
What Post-Production Scoring Can't Do
A score on a production trace arrives after a customer has had the experience. Three gaps follow from that timing:
- It can't stop a release. By the time a trace is scored, the version that produced it is live.
- It grades the account, not the effect. A trace records what the agent said and which tools it called. It does not, by default, check what changed in the file, the ticket, or the database. An agent's own summary is the weakest evidence of what it did.
- It only sees the traffic you happened to get. Prompt injection, tool misuse, and rare policy edge cases show up in production only after someone tries them.
Each gap closes only with AI agent testing before production. The breakdown of LLM evaluation vs agent testing shows where a model score and a build gate disagree.
What Pre-Production Agent Testing Looks Like
Agent Assurance tests and grades what an AI agent actually changed, not what it says it did. It is delivered by Rook CLI, runs before release, and works in four phases:
- Discover. Point Rook at a repository, a PRD, or a folder of policy docs. It reads the agent's declared tools, including tools exposed through MCP servers. Anything the agent doesn't declare comes back as unknown, never guessed.
- Generate. Rook derives functional and adversarial scenarios, each with its own gradable criteria. Prompt injection, instruction override, and tool misuse are generated by default. Nobody writes tests by hand.
- Run and judge. The agent is invoked for real. Judges watch files, artifacts, and every tool call, check calls against the declared tool surface, and change nothing themselves.
- Report. Each criterion gets one of three verdicts: pass, fail, or unable to verify. Unverifiable criteria are reported separately and never counted in the pass rate. That separate figure is the assurance gap.
A real example from the product: an agent replied that a Discord bot wasn't configured and the message could not be sent. The tool call Rook recorded showed the message was sent. Verdict: fail.
A trace-scoring judge reading only the reply would have graded the agent's honesty about a failure; Rook graded the effect.
Introducing Agent Assurance explains why the suite comes from the agent's codebase rather than from hand-written tests, and the guide to multi-agent testing applies the same effect-first grading when agents hand work to each other.
For agents judged on what they say rather than what they change, TestMu AI covers chat, voice and IVR agent testing as a separate product.
Pre-Production Agent Testing vs Post-Production at a Glance
The table compares pre-release agent evals vs observability, row by row.
| Pre-production (Agent Assurance) | Post-production (observability evals) | |
|---|---|---|
| Question it answers | Should this version ship? | How is the shipped version behaving? |
| When it runs | Before release, in CI | On live traffic, after release |
| Test inputs | Scenarios generated from code, PRD or policy docs | Production traces, plus datasets you build for pre-release experiments |
| What gets graded | What the run changed: files, artifacts, tool calls | What the agent said and recorded in the trace |
| Adversarial cases | Generated by default | From datasets you supply, or opt-in red-teaming |
| What it can block | A merge or a deploy | An individual output, via runtime guardrails |
| Unverifiable results | Reported as the assurance gap | No separate "could not verify" verdict |
| Owner | Engineers building agents, QA, platform teams | SRE, operations, AI platform teams |
How the Two Work Together
The strongest setup runs both, in a loop:
- Every agent change runs through Agent Assurance in CI. The build passes only on per-criterion verdicts and an acceptable assurance gap.
- The release ships. Your observability platform scores it on live traffic and guards it at runtime.
- When production surfaces a new failure, write it down as a requirement or a policy line in the spec Rook reads. The next generated suite tests for it before the following release.
Step 3 is agent regression testing: every production failure becomes a scenario the next release has to pass.

Pre-production stops known failures from shipping. Post-production finds the unknown ones. Each feeds the other.
Gate on the Report, Not the Exit Code
One practical warning for CI. Rook's exit code 0 means the run finished, not that the agent passed; every scenario within it could have failed. Exit code 1 means the command itself failed.
Gate the build in two steps: confirm the run covered every selected scenario, then fail on the per-criterion verdicts. TestMu AI publishes gate recipes for GitHub Actions, Jenkins, and Argo CD.
Try It on Your Agent
Rook is publicly installable on macOS, Linux, Windows x64, or WSL, without Docker.
brew install lambdatest/rook/rook
# or
npm install -g @testmuai/rookPrefer to drive it from a coding agent? npx @testmuai/rook-skill adds a /rook skill to Claude Code, Codex CLI, Gemini CLI, and others. Writes are real, so point your first run at staging.
Kane CLI is its sibling for coding-agent users: it checks the web app an agent builds in a real browser.
Start with the Agent Assurance quickstart, or book a demo to see it gate an enterprise agent pipeline.
Sources
- Dynatrace completes acquisition of Arize
- Dynatrace: agreement to acquire Arize ($915M)
- Splunk Observability at .conf26 (Agent Observability, powered by Galileo)
- Splunk: Cisco completes Galileo acquisition
- New Relic: Introducing AI Evaluation (Oct 6, 2026)
- New Relic: AI Evaluation press release (public preview in November)
- Arize AX: CI/CD with experiments
- Galileo: Experiments basics (CI/CD)
- TestMu AI Agent Assurance product page
Author
Vipul Verma is Group Senior Vice President of Engineering at TestMu AI (formerly LambdaTest), where he heads the entire engineering organization that builds KaneAI, HyperExecute, and the broader testing cloud. He brings 15+ years architecting, securing, and scaling large enterprise applications across multiple sites. Before TestMu AI he was India Head at LogicHub, where he built the India R&D site from the first employee to a 30-plus engineering team, and Principal Software Engineer at Sumo Logic, where he was the first engineer in the India office and shipped search-performance and pricing-model initiatives. Earlier he worked on trading platforms at Portware and D. E. Shaw. Vipul holds a B.Tech in Computer Science from IIT Kharagpur.
Reviewer
Mayank Bhola is Co-Founder and Head of Products at TestMu AI (formerly LambdaTest), where he leads the entire product portfolio across KaneAI, Kane CLI, HyperExecute, SmartUI, the Real Device Cloud, Accessibility, and other software testing product lines. As an early Lead Architect he designed and built the company's flagship Tunnel technology from scratch, created the React-based automation platform, and architected the data-intensive pipelines and FAAS services that scale it. He brings more than 10 years of experience in software development and product engineering, with earlier roles as Head of Technology at Juggernaut Books and Senior Software Engineer at PressPlay TV and Zomato. Mayank holds a B.Tech in Computer Engineering from JIIT Noida.
Pre-Production Agent Testing FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests



