Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Learning Hub
- /
- Agent Regression Testing: Catch What Breaks Between Releases
Agent Regression Testing: Catch What Breaks Between Releases
Agent regression testing guide: what breaks AI agents between releases, what to assert, how to tell a regression from a flaky run, and how to gate on verdicts.
Published on:
OVERVIEW
A developer rewords one sentence in the system prompt of an agent that files support escalations, and the agent's replies read as clearly as before. The change shows up only in the tickets: the agent has stopped attaching the customer's account ID, so each escalation lands in a queue nobody watches. A check that grades the reply cannot see that, and catching it before release is the job of agent regression testing.
Prompt edits are one trigger among several. In a July 2026 study of 35 Qwen Code CLI releases, Ben Sghaier, Li, Adams and Hassan held the model constant and changed only the harness around it. The agent's resolve rate on 50 SWE-bench Verified tasks peaked at 39.0% in some of the earliest versions and dipped as low as 23.0% in some mid-cycle updates.
Overview
Agent regression testing reruns a fixed suite of scenarios against an AI agent after any change that could alter its behavior, such as a prompt edit, model upgrade, tool schema change or code release, and compares each scenario's verdict with the previous run. A scenario that passed before the change and fails after it is a regression.
What Does an Agent Regression Test Compare?
- Newly failing scenarios: A scenario that passed on the previous run and fails now, with nothing about the scenario changed, is the regression signal. Hold the release until someone confirms the new behavior was intended, because the agent's reply can still read correctly while the action behind it fails.
- Newly fixed scenarios: A scenario that failed on the previous run and passes now suggests a change worked; confirm your change is the reason. Keep it in the suite after the fix, because from then on it is the check that stops the same failure from returning unnoticed.
- Flaky scenarios: A flaky scenario flips between pass and fail across runs while nothing about it changed. It needs more repeats or a more precise criterion, the opposite of the rollback a real regression calls for, so flaky results stay out of the regression count.
- Changed definitions: When a scenario's goal or criteria were edited, its earlier verdicts no longer apply to the new run. Treat the new result as that scenario's first baseline instead of reading it as a regression or a fix.
- Unverified criteria: A criterion that no evidence from the run could check is neither a pass nor a failure. TestMu AI's Agent Assurance reports these as Unable to Verify and keeps them out of the pass rate, so an unverified criterion never counts for or against the agent.
What Is Agent Regression Testing?
Agent regression testing is the practice of checking an AI agent for behavior that got worse after a change, by rerunning the same scenarios on the new version and comparing each verdict with the last run. The change can be anything the agent depends on, such as a prompt edit or a model upgrade, and a regression can be a wrong tool call as easily as a record that was never written.
The idea comes from classic regression testing, which reruns existing tests after a code change to confirm that working features still work. Microsoft's public AI agent evaluation scenario library on GitHub describes the agent version as scenarios for "validating that agent updates don't break existing behavior", to run before publishing a knowledge source update, a topic change, a tool configuration change or a prompt adjustment.
These terms sound alike and describe different work:
| Term | What is under test | What a regression looks like |
|---|---|---|
| Agent regression testing | An AI agent that acts, across its own versions | A scenario that passed on the last run fails after a change to the agent's prompt, model, tools, code or data |
| Agentic regression testing | An ordinary application, with AI agents selecting, running and repairing its regression tests | An application bug that slips through because an agent skipped the test or repaired it until it passed |
| LLM regression testing | A model's outputs across model versions | Items that flip from right to wrong after a model upgrade, even when the average score improves |
| Classic regression testing | Deterministic application code | An assertion that passed before a code change fails after it |
This guide covers the first row, for agents that act: they call tools, write files and change records. The guide to agentic regression testing covers the second row, where agents run an application's tests, and LLM regression testing covers model upgrades measured on their own.
An agent that talks to people over chat, voice or phone regresses in what it says, which TestMu AI grades with its separate Agent Testing product. The guide to AI voice agent regression testing covers the voice case.
Agent Regression Testing vs Traditional Regression Testing
A traditional regression suite can assert an exact result because the code under test is deterministic: the same input gives the same output, so a failure usually reproduces. An agent breaks that assumption in the places a suite depends on:
- Output - the same request can produce different wording, different tool calls and a different path, and more than one of them can be correct.
- Evidence - an agent's result is often an action, such as a ticket filed or a file written, so the test has to check the effect rather than a returned value.
- Change surface - prompts, tool definitions, knowledge sources, the harness and a hosted model can each change the agent's behavior, and knowledge sources or a hosted model can change without any release of your own.
- Verdicts - some criteria cannot be checked with the evidence a run leaves, so a result can be undecided as well as passed or failed.
Variance alone can swamp real differences. In Identical Runs, Different Results (September 2026), Ariño de la Rubia and Pafka ran six agent-model pairings 52 times each under fixed settings on one machine-learning task, and "identical runs of one pairing varied more than the pairings differed from one another".
Run each scenario more than once, or the suite will report regressions that are noise and miss regressions that are real.
Which Changes Cause Agent Regressions?
Any change to the prompt, the model, the tools, the harness code or the data the agent reads can cause a regression. Map each kind of change to the behavior it tends to break and the evidence that shows it, and the rerun set follows:
| Change | What tends to regress | Evidence that shows it | What to rerun |
|---|---|---|---|
| System prompt or instruction file | Which tool the agent picks, what it refuses and which fields it fills in | The tool calls the run made and the records or files it created | Scenarios that exercise the edited instruction, plus the adversarial scenarios |
| Model or model version | Tool choice, argument format, refusals, token use and latency | Tool calls, token usage and latency per scenario | The full suite, with repeats |
| Tool schema or tool description | Whether the agent calls the tool, and the arguments it passes | Observed calls checked against the tools the agent declares | Scenarios that use the changed tool, plus its must-not-call rules |
| MCP server or external API | Calls that now error, return a new shape or change the wrong record | Tool results, errors and a read-only check of the record | Integration scenarios for that server or API |
| Harness or orchestration code | Loop limits, context handling and retries, plus tokens and tool calls per task | Turns, tool-call counts and token usage against budgets | The full suite, with cost and latency budgets |
| Knowledge sources | Answers and actions that depend on the changed documents | The reply, plus the record or file the action leaves | Scenarios that read the changed sources, plus a random sample of the rest |
The harness row is the easiest to miss: the Qwen Code study's authors note that practitioners regularly report quality regressions after harness updates, yet attribute them to the underlying model. Token use in the same study grew from about 391K per task in the first nine releases to nearly 668K in the latest ones, an increase of over 70% with no corresponding improvement in resolve rate. A pass-fail check misses that kind of regression; a token or tool-call budget per scenario catches it.
Knowledge sources need a wider net than the changed documents alone. Microsoft's scenario library warns that knowledge source changes have a blast radius that is often wider than expected, which is why the rerun set for a knowledge update includes a random sample of unrelated scenarios.
What Should an Agent Regression Test Assert?
Assert the criteria the outcome must meet, checked against evidence the agent did not write, and leave the agent free to choose its own path.
- Outcome criteria - state what must be true after the run, such as the escalation ticket sitting in the right queue with the account ID attached. Grade each criterion on its own, so a run that meets three criteria and misses a fourth shows exactly which one broke.
- Effects over the agent's own report - check the record, file or artifact the run left, because the agent's summary of its own work is a claim. A read-only query of the ticket system settles whether the ticket exists.
- Must-not-call rules - list the tools a scenario must never call, such as a delete or a payment, and grade them from the calls the run actually made.
- Budgets instead of exact counts - set ceilings for tokens, turns and tool calls per scenario. An exact count fails on normal variation, while a ceiling still catches the harness-style growth described above.
- Compliance beside quality - grade the task's rules as their own criteria, separate from any quality score.
The last point is where score-only suites go wrong. In the Identical Runs study, fewer than one run in twenty broke the task's data rules, but those runs held the highest scores, and the authors recommend reporting compliance beside quality.
Avoid asserting the exact sequence of tool calls unless one call needs the output of another. Two correct runs can reach the same outcome through different calls, so a sequence assertion fails runs that did the right thing. Microsoft's library re-verifies "that the correct topics, flows, and actions still fire for previously tested inputs", and checks order where the design sets it: the steps of a modified conversation flow, and a chained tool workflow in which one tool's output feeds the next. Whether an action's effect landed in the target system is a separate criterion.
How to Separate a Regression From Noise
Repeat each scenario, compare how often it passes on the baseline, the last release, with how often it passes on the candidate, the version with the change, and call it a regression only when the difference is consistent.
Precise comparisons need many runs: resolving the differences between agents in the Identical Runs study would take "tens to more than a hundred runs of each", in the authors' words. A regression suite can settle for fewer because it looks for scenarios that stopped working, not for a few points of score.
- Fix the repeat count first - decide how many times each scenario runs before looking at any result, use more repeats for scenarios with open-ended paths, and keep the count the same on the baseline and the candidate.
- Compare pass counts - a scenario that passed every repeat on the baseline and fails most or all repeats on the candidate is a regression, and one that fails some repeats on both is flaky. A partial drop on the candidate alone needs more repeats before anyone decides.
- Hold everything else still - run the baseline and the candidate against the same staging data, tool versions and invocation settings, so the change under test is the only difference.
- Keep the baseline with the release - store each release's verdicts, so the next change compares against what shipped rather than against an old run.
How to Read a Run-Over-Run Diff
Put every scenario into one state before anyone starts debugging, because each state calls for a different response:
| State | What it means | What to do |
|---|---|---|
| Newly failing | Passed on the baseline, fails on the candidate, scenario unchanged | Block the release, read the failed criteria and their evidence, then confirm whether the new behavior was intended |
| Newly fixed | Failed on the baseline, passes on the candidate | Confirm your change is the reason, then keep the scenario as the guard for that fix |
| Flaky | Flips between runs while the scenario is unchanged | Do not count it as a regression; look for what varies between runs, such as a time-dependent fixture or a criterion the judge can read two ways |
| Changed definition | The scenario's goal or criteria were edited | Treat the new result as its first baseline, since the old verdicts measured a different scenario |
| Still failing | Failed on both runs | Track it as a known issue, apart from new failures, so it cannot hide one |
| Unverified | No evidence from the run could check a criterion | Keep it out of the pass rate and add the evidence source that would close the gap |
Microsoft's regression scenarios ask for a similar triage of every new failure: true regressions, expected changes and flaky tests, the last defined as "intermittent failures unrelated to the change". An expected change means the scenario itself is out of date, so update it, and it moves to the changed-definition state for its next run.
How to Gate a Release on Regression Verdicts
Gate on the per-scenario verdicts in the test report instead of the runner's exit code alone, and decide before the first run how the gate treats results that could not be verified.
- Verdicts over exit codes - with some agent test runners, the exit code tells you only whether the run finished: they exit 0 on a finished run whatever its scenarios found, so check your runner's documentation.
- Two numbers, two denominators - report the pass rate over scenarios with a decided verdict and, beside it, the share of criteria the run could verify. A high pass rate over little verified evidence is weak release evidence.
- An unverified-results policy - a strict gate blocks on any unverified result, and a lenient one reports them next to the pass rate. Pick one before the first run, and never count an unverified result as a pass or a failure.
- Block on new failures - fail the gate on newly failing scenarios, and let a still-failing one through only when someone has explicitly accepted and documented it, so the gate is not red on every run. The platform examples in TestMu AI's CI/CD guide are stricter: every selected scenario must pass.
- Thresholds from a baseline - measure the pass rate and verification coverage over a few releases before you set thresholds on either.
- A reviewed branch and staging - an agent under test makes real writes, so run the gate from a reviewed branch against a staging deployment with its own credentials.
Beside the new-failure rule, Microsoft's regression scenarios suggest a pass-rate threshold for publishing, such as 95% of regression test cases passing, and warn that when more than 10% fail after a typical knowledge update, the suite or the knowledge architecture may need restructuring.
How to Grow the Suite From Production Failures
Turn every agent failure that AI agent monitoring finds after release into a scenario before you fix it, so the fix gets a test that proves it and the failure cannot return unnoticed.
- Capture the input and the context - save the user turn, the documents and tool results the agent saw, and the state of the records it touched.
- Write the broken criterion - state what should have been true as an effect you can check, such as a field on the ticket or a tool that must not be called.
- Confirm it fails first - run the new scenario against the current release. A scenario that passes before the fix proves nothing about the fix.
- Fix, rerun and keep it - after the fix the scenario shows as newly fixed, and from then on it guards that behavior.
- Retire without deleting - when a scenario goes stale, leave it out of runs but keep it and its history, so it can come back.
For coverage beyond escaped failures, Microsoft's library suggests at least the top 3 to 5 inputs for every topic, plus a random sample of 10 to 15 general inputs to catch unexpected routing changes.
A suite that keeps growing gets slow to rerun. She and Lin's study of efficient benchmarking for an evolving production agent used 574 historical runs of an analytics agent's benchmark and found that adaptive testing on 200 questions, 38.5% of a full run, came within 1.03 percentage points of the full-run score on average. The team deployed difficulty-stratified fixed subsets instead, for their operational simplicity.
Run a stratified subset on each change and the full suite before release. A subset estimates the overall score well and can still miss the one scenario that regressed, which is why the full run stays as the release gate.
Running Agent Regression Tests With Agent Assurance
TestMu AI's Agent Assurance, which runs from the terminal as Rook CLI, applies this practice to agents that act. It derives the scenario suite from the agent's code, or from a PRD or spec when there is no code to read, invokes the real agent and grades each criterion against evidence. Rerun the suite after a change, and the run reports what changed since the last one.
For a regression run, Agent Assurance checks:
- Run-over-run change - newly failing, fixed and flaky scenarios are reported apart, and scenarios whose definition changed are marked because their history no longer compares. The Rook CLI README defines flaky as a scenario that flips between runs while unchanged.
- Tool calls against declared tools - the calls a run makes are checked against the tools the agent itself declares, and must-not-call criteria are graded from observed calls when the profile returns them.
- Effects the run leaves - files that changed on disk under the paths the profile declares, artifacts produced, and records confirmed with a read-only query through a tool you approve (stdio MCP servers only in the current release).
- Unable to Verify - a criterion no evidence could check is reported as unverifiable, never as a pass or a failure, and is excluded from the pass rate's denominator.
- Verdicts for the gate - a finished run exits 0 whether its scenarios passed or failed, so a CI gate reads the verdicts in
rook report --json. The walkthrough on how to gate GitHub Actions on Rook CLI verdicts builds that gate step by step.
Point regression runs at staging, because the agent's writes are real.
Between releases, you curate the suite from the terminal. This is rook scenarios --help, captured from Rook CLI 0.1.5 on 30 September 2026:
Usage: rook scenarios [options] [command]
inspect and curate the test set
Options:
-h, --help display help for command
Commands:
list [options] what would run, and what could not
exclude [options] <ids...> keep a scenario on disk but leave it out of runs
include [options] <ids...> undo an exclude
delete [options] <ids...> remove scenarios permanently
help [command] display help for commandFor a regression suite, exclude is the safer way to retire a scenario: it stays on disk and out of runs, and include brings it back once its flakiness is fixed. rook scenarios list shows what would run, and what could not, before a run starts.
Rook CLI flags a scenario as flaky from its history across runs. Within a run, the scenario's repeat field sets how many samples it takes, and the Agent Assurance scenarios guide advises raising it only where several samples answer a real reliability or performance question.
Most eval and observability tools score what your agent said and recorded. Agent Assurance checks what the run changed, and reports what it could not verify. Eval tools rerun cases you wrote or synthesized from documents and traces, while Agent Assurance derives its scenarios from the agent's code; like eval tools, it uses model judges, grades what the agent says as well as what it did, and runs in CI.
| Agent Assurance | AI eval tools | LLM observability | |
|---|---|---|---|
| Test cases | From your code or spec | Written or synthesized | From production traces |
| Tool calls | Against declared tools | Against your lists | Logged, optionally scored |
| Side effects | Files, artifacts, probes | Scripted per task | Trace data only |
| Grading | Claimed actions aren't proof | LLM judge or code checks | LLM judge or human review |
| Adversarial tests | Generated by default | Add-on in some tools | Not generated |
| When it runs | Before release, in CI | CI and live traffic | Production, plus CI |
| Unverifiable results | Reported separately | Errors or opt-in skips | Left unscored |
The eval and observability columns describe each category's default approach, not any single product.
Rook CLI installs from npm on Node.js 22 or newer, and the Rook CLI installation guide covers Homebrew and the shell installer, which bring their own Node runtime:
npm install -g @testmuai/rook
rook --versionTo drive it from Claude Code, install the skill, then ask for the regression run in a /rook request:
npx @testmuai/rook-skill@latest install --agent claude-code/rook Rerun this agent's current scenarios on the staging profile as a regression run. Before invoking the agent, show me the plan and which writes it could make in staging. After the run, compare it with the previous run: list newly failing, newly fixed and flaky scenarios, and any whose definition changed, with each failed criterion and its evidence.Note: Agent Assurance keeps scenarios, runs and evidence as plain, committable files in your project, so each release's baseline can sit next to the code it tested. Check the evidence for secrets and personal data before you commit it.
Conclusion
Start agent regression testing with the change your team ships most often. Pick the scenarios that exercise it hardest, run each several times on the current release to record a baseline, and rerun them before the next change ships.
Sort the results into newly failing, newly fixed, flaky, changed, still failing and unverified, and let newly failing scenarios, plus any still-failing one nobody has accepted, block the release. If your agent runs on an open-source harness, expect changes you did not make: Ben Sghaier and colleagues counted OpenCode at 18.0 releases per week.
Then move the same gate into your pipeline, following the guide to run Agent Assurance in CI/CD, with every run pointed at staging.
Author
Samyak Goyal is a Senior Member of Technical Staff at TestMu AI engineering Kane CLI, the command-line tool that runs browser automation from the terminal, where a flow described in natural language executes in a real Chrome browser and returns pass or fail with shareable proof. He is a backend engineer with 4+ years of experience, previously an SDE at Innovaccer, where he built APIs, introduced Kafka, and cut deployment from weeks to hours. Samyak also builds multi-agent systems, skill-orchestration frameworks, and a personal copilot that indexes 200+ microservice repositories.
Reviewer
Srinivasan Sekar is Director of Engineering at TestMu AI (formerly LambdaTest), where he leads engineering and open-source initiatives behind the Selenium and Appium automation grid and owns TestMu AI's MCP Server. A committer to Appium and a contributor to Selenium, WebdriverIO, Taiko, and AppiumTestDistribution, he brings over 15 years of experience in quality engineering and open-source technologies. He is the author of the Apress book 'The MCP Standard: A Developer's Guide to Building Universal AI Tools with the Model Context Protocol,' a Certified Kubernetes and Cloud Native Associate, and an international conference speaker. Before TestMu AI he spent over eight years at Thoughtworks as a Principal Consultant and Quality Architect. Srinivasan holds a B.Tech in Information Technology from Anna University.
Agent Regression Testing FAQs
Did you find this page helpful?
More Related Learning Hubs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests



