Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIAI TestingRegression Testing

Agent Regression Testing: Catch What Breaks Between Releases

Agent regression testing guide: what breaks AI agents between releases, what to assert, how to tell a regression from a flaky run, and how to gate on verdicts.

Published on:

OVERVIEW

A developer rewords one sentence in the system prompt of an agent that files support escalations, and the agent's replies read as clearly as before. The change shows up only in the tickets: the agent has stopped attaching the customer's account ID, so each escalation lands in a queue nobody watches. A check that grades the reply cannot see that, and catching it before release is the job of agent regression testing.

Prompt edits are one trigger among several. In a July 2026 study of 35 Qwen Code CLI releases, Ben Sghaier, Li, Adams and Hassan held the model constant and changed only the harness around it. The agent's resolve rate on 50 SWE-bench Verified tasks peaked at 39.0% in some of the earliest versions and dipped as low as 23.0% in some mid-cycle updates.

Overview

Agent regression testing reruns a fixed suite of scenarios against an AI agent after any change that could alter its behavior, such as a prompt edit, model upgrade, tool schema change or code release, and compares each scenario's verdict with the previous run. A scenario that passed before the change and fails after it is a regression.

What Does an Agent Regression Test Compare?

  • Newly failing scenarios: A scenario that passed on the previous run and fails now, with nothing about the scenario changed, is the regression signal. Hold the release until someone confirms the new behavior was intended, because the agent's reply can still read correctly while the action behind it fails.
  • Newly fixed scenarios: A scenario that failed on the previous run and passes now suggests a change worked; confirm your change is the reason. Keep it in the suite after the fix, because from then on it is the check that stops the same failure from returning unnoticed.
  • Flaky scenarios: A flaky scenario flips between pass and fail across runs while nothing about it changed. It needs more repeats or a more precise criterion, the opposite of the rollback a real regression calls for, so flaky results stay out of the regression count.
  • Changed definitions: When a scenario's goal or criteria were edited, its earlier verdicts no longer apply to the new run. Treat the new result as that scenario's first baseline instead of reading it as a regression or a fix.
  • Unverified criteria: A criterion that no evidence from the run could check is neither a pass nor a failure. TestMu AI's Agent Assurance reports these as Unable to Verify and keeps them out of the pass rate, so an unverified criterion never counts for or against the agent.

What Is Agent Regression Testing?

Agent regression testing is the practice of checking an AI agent for behavior that got worse after a change, by rerunning the same scenarios on the new version and comparing each verdict with the last run. The change can be anything the agent depends on, such as a prompt edit or a model upgrade, and a regression can be a wrong tool call as easily as a record that was never written.

The idea comes from classic regression testing, which reruns existing tests after a code change to confirm that working features still work. Microsoft's public AI agent evaluation scenario library on GitHub describes the agent version as scenarios for "validating that agent updates don't break existing behavior", to run before publishing a knowledge source update, a topic change, a tool configuration change or a prompt adjustment.

These terms sound alike and describe different work:

TermWhat is under testWhat a regression looks like
Agent regression testingAn AI agent that acts, across its own versionsA scenario that passed on the last run fails after a change to the agent's prompt, model, tools, code or data
Agentic regression testingAn ordinary application, with AI agents selecting, running and repairing its regression testsAn application bug that slips through because an agent skipped the test or repaired it until it passed
LLM regression testingA model's outputs across model versionsItems that flip from right to wrong after a model upgrade, even when the average score improves
Classic regression testingDeterministic application codeAn assertion that passed before a code change fails after it

This guide covers the first row, for agents that act: they call tools, write files and change records. The guide to agentic regression testing covers the second row, where agents run an application's tests, and LLM regression testing covers model upgrades measured on their own.

An agent that talks to people over chat, voice or phone regresses in what it says, which TestMu AI grades with its separate Agent Testing product. The guide to AI voice agent regression testing covers the voice case.

Agent Regression Testing vs Traditional Regression Testing

A traditional regression suite can assert an exact result because the code under test is deterministic: the same input gives the same output, so a failure usually reproduces. An agent breaks that assumption in the places a suite depends on:

  • Output - the same request can produce different wording, different tool calls and a different path, and more than one of them can be correct.
  • Evidence - an agent's result is often an action, such as a ticket filed or a file written, so the test has to check the effect rather than a returned value.
  • Change surface - prompts, tool definitions, knowledge sources, the harness and a hosted model can each change the agent's behavior, and knowledge sources or a hosted model can change without any release of your own.
  • Verdicts - some criteria cannot be checked with the evidence a run leaves, so a result can be undecided as well as passed or failed.

Variance alone can swamp real differences. In Identical Runs, Different Results (September 2026), Ariño de la Rubia and Pafka ran six agent-model pairings 52 times each under fixed settings on one machine-learning task, and "identical runs of one pairing varied more than the pairings differed from one another".

Run each scenario more than once, or the suite will report regressions that are noise and miss regressions that are real.

Which Changes Cause Agent Regressions?

Any change to the prompt, the model, the tools, the harness code or the data the agent reads can cause a regression. Map each kind of change to the behavior it tends to break and the evidence that shows it, and the rerun set follows:

ChangeWhat tends to regressEvidence that shows itWhat to rerun
System prompt or instruction fileWhich tool the agent picks, what it refuses and which fields it fills inThe tool calls the run made and the records or files it createdScenarios that exercise the edited instruction, plus the adversarial scenarios
Model or model versionTool choice, argument format, refusals, token use and latencyTool calls, token usage and latency per scenarioThe full suite, with repeats
Tool schema or tool descriptionWhether the agent calls the tool, and the arguments it passesObserved calls checked against the tools the agent declaresScenarios that use the changed tool, plus its must-not-call rules
MCP server or external APICalls that now error, return a new shape or change the wrong recordTool results, errors and a read-only check of the recordIntegration scenarios for that server or API
Harness or orchestration codeLoop limits, context handling and retries, plus tokens and tool calls per taskTurns, tool-call counts and token usage against budgetsThe full suite, with cost and latency budgets
Knowledge sourcesAnswers and actions that depend on the changed documentsThe reply, plus the record or file the action leavesScenarios that read the changed sources, plus a random sample of the rest

The harness row is the easiest to miss: the Qwen Code study's authors note that practitioners regularly report quality regressions after harness updates, yet attribute them to the underlying model. Token use in the same study grew from about 391K per task in the first nine releases to nearly 668K in the latest ones, an increase of over 70% with no corresponding improvement in resolve rate. A pass-fail check misses that kind of regression; a token or tool-call budget per scenario catches it.

Knowledge sources need a wider net than the changed documents alone. Microsoft's scenario library warns that knowledge source changes have a blast radius that is often wider than expected, which is why the rerun set for a knowledge update includes a random sample of unrelated scenarios.

What Should an Agent Regression Test Assert?

Assert the criteria the outcome must meet, checked against evidence the agent did not write, and leave the agent free to choose its own path.

  • Outcome criteria - state what must be true after the run, such as the escalation ticket sitting in the right queue with the account ID attached. Grade each criterion on its own, so a run that meets three criteria and misses a fourth shows exactly which one broke.
  • Effects over the agent's own report - check the record, file or artifact the run left, because the agent's summary of its own work is a claim. A read-only query of the ticket system settles whether the ticket exists.
  • Must-not-call rules - list the tools a scenario must never call, such as a delete or a payment, and grade them from the calls the run actually made.
  • Budgets instead of exact counts - set ceilings for tokens, turns and tool calls per scenario. An exact count fails on normal variation, while a ceiling still catches the harness-style growth described above.
  • Compliance beside quality - grade the task's rules as their own criteria, separate from any quality score.

The last point is where score-only suites go wrong. In the Identical Runs study, fewer than one run in twenty broke the task's data rules, but those runs held the highest scores, and the authors recommend reporting compliance beside quality.

Avoid asserting the exact sequence of tool calls unless one call needs the output of another. Two correct runs can reach the same outcome through different calls, so a sequence assertion fails runs that did the right thing. Microsoft's library re-verifies "that the correct topics, flows, and actions still fire for previously tested inputs", and checks order where the design sets it: the steps of a modified conversation flow, and a chained tool workflow in which one tool's output feeds the next. Whether an action's effect landed in the target system is a separate criterion.

How to Separate a Regression From Noise

Repeat each scenario, compare how often it passes on the baseline, the last release, with how often it passes on the candidate, the version with the change, and call it a regression only when the difference is consistent.

Precise comparisons need many runs: resolving the differences between agents in the Identical Runs study would take "tens to more than a hundred runs of each", in the authors' words. A regression suite can settle for fewer because it looks for scenarios that stopped working, not for a few points of score.

  • Fix the repeat count first - decide how many times each scenario runs before looking at any result, use more repeats for scenarios with open-ended paths, and keep the count the same on the baseline and the candidate.
  • Compare pass counts - a scenario that passed every repeat on the baseline and fails most or all repeats on the candidate is a regression, and one that fails some repeats on both is flaky. A partial drop on the candidate alone needs more repeats before anyone decides.
  • Hold everything else still - run the baseline and the candidate against the same staging data, tool versions and invocation settings, so the change under test is the only difference.
  • Keep the baseline with the release - store each release's verdicts, so the next change compares against what shipped rather than against an old run.

How to Read a Run-Over-Run Diff

Put every scenario into one state before anyone starts debugging, because each state calls for a different response:

StateWhat it meansWhat to do
Newly failingPassed on the baseline, fails on the candidate, scenario unchangedBlock the release, read the failed criteria and their evidence, then confirm whether the new behavior was intended
Newly fixedFailed on the baseline, passes on the candidateConfirm your change is the reason, then keep the scenario as the guard for that fix
FlakyFlips between runs while the scenario is unchangedDo not count it as a regression; look for what varies between runs, such as a time-dependent fixture or a criterion the judge can read two ways
Changed definitionThe scenario's goal or criteria were editedTreat the new result as its first baseline, since the old verdicts measured a different scenario
Still failingFailed on both runsTrack it as a known issue, apart from new failures, so it cannot hide one
UnverifiedNo evidence from the run could check a criterionKeep it out of the pass rate and add the evidence source that would close the gap

Microsoft's regression scenarios ask for a similar triage of every new failure: true regressions, expected changes and flaky tests, the last defined as "intermittent failures unrelated to the change". An expected change means the scenario itself is out of date, so update it, and it moves to the changed-definition state for its next run.

How to Gate a Release on Regression Verdicts

Gate on the per-scenario verdicts in the test report instead of the runner's exit code alone, and decide before the first run how the gate treats results that could not be verified.

  • Verdicts over exit codes - with some agent test runners, the exit code tells you only whether the run finished: they exit 0 on a finished run whatever its scenarios found, so check your runner's documentation.
  • Two numbers, two denominators - report the pass rate over scenarios with a decided verdict and, beside it, the share of criteria the run could verify. A high pass rate over little verified evidence is weak release evidence.
  • An unverified-results policy - a strict gate blocks on any unverified result, and a lenient one reports them next to the pass rate. Pick one before the first run, and never count an unverified result as a pass or a failure.
  • Block on new failures - fail the gate on newly failing scenarios, and let a still-failing one through only when someone has explicitly accepted and documented it, so the gate is not red on every run. The platform examples in TestMu AI's CI/CD guide are stricter: every selected scenario must pass.
  • Thresholds from a baseline - measure the pass rate and verification coverage over a few releases before you set thresholds on either.
  • A reviewed branch and staging - an agent under test makes real writes, so run the gate from a reviewed branch against a staging deployment with its own credentials.

Beside the new-failure rule, Microsoft's regression scenarios suggest a pass-rate threshold for publishing, such as 95% of regression test cases passing, and warn that when more than 10% fail after a typical knowledge update, the suite or the knowledge architecture may need restructuring.

How to Grow the Suite From Production Failures

Turn every agent failure that AI agent monitoring finds after release into a scenario before you fix it, so the fix gets a test that proves it and the failure cannot return unnoticed.

  • Capture the input and the context - save the user turn, the documents and tool results the agent saw, and the state of the records it touched.
  • Write the broken criterion - state what should have been true as an effect you can check, such as a field on the ticket or a tool that must not be called.
  • Confirm it fails first - run the new scenario against the current release. A scenario that passes before the fix proves nothing about the fix.
  • Fix, rerun and keep it - after the fix the scenario shows as newly fixed, and from then on it guards that behavior.
  • Retire without deleting - when a scenario goes stale, leave it out of runs but keep it and its history, so it can come back.

For coverage beyond escaped failures, Microsoft's library suggests at least the top 3 to 5 inputs for every topic, plus a random sample of 10 to 15 general inputs to catch unexpected routing changes.

A suite that keeps growing gets slow to rerun. She and Lin's study of efficient benchmarking for an evolving production agent used 574 historical runs of an analytics agent's benchmark and found that adaptive testing on 200 questions, 38.5% of a full run, came within 1.03 percentage points of the full-run score on average. The team deployed difficulty-stratified fixed subsets instead, for their operational simplicity.

Run a stratified subset on each change and the full suite before release. A subset estimates the overall score well and can still miss the one scenario that regressed, which is why the full run stays as the release gate.

Running Agent Regression Tests With Agent Assurance

TestMu AI's Agent Assurance, which runs from the terminal as Rook CLI, applies this practice to agents that act. It derives the scenario suite from the agent's code, or from a PRD or spec when there is no code to read, invokes the real agent and grades each criterion against evidence. Rerun the suite after a change, and the run reports what changed since the last one.

For a regression run, Agent Assurance checks:

  • Run-over-run change - newly failing, fixed and flaky scenarios are reported apart, and scenarios whose definition changed are marked because their history no longer compares. The Rook CLI README defines flaky as a scenario that flips between runs while unchanged.
  • Tool calls against declared tools - the calls a run makes are checked against the tools the agent itself declares, and must-not-call criteria are graded from observed calls when the profile returns them.
  • Effects the run leaves - files that changed on disk under the paths the profile declares, artifacts produced, and records confirmed with a read-only query through a tool you approve (stdio MCP servers only in the current release).
  • Unable to Verify - a criterion no evidence could check is reported as unverifiable, never as a pass or a failure, and is excluded from the pass rate's denominator.
  • Verdicts for the gate - a finished run exits 0 whether its scenarios passed or failed, so a CI gate reads the verdicts in rook report --json. The walkthrough on how to gate GitHub Actions on Rook CLI verdicts builds that gate step by step.

Point regression runs at staging, because the agent's writes are real.

Between releases, you curate the suite from the terminal. This is rook scenarios --help, captured from Rook CLI 0.1.5 on 30 September 2026:

Usage: rook scenarios [options] [command]

inspect and curate the test set

Options:
  -h, --help                  display help for command

Commands:
  list [options]              what would run, and what could not
  exclude [options] <ids...>  keep a scenario on disk but leave it out of runs
  include [options] <ids...>  undo an exclude
  delete [options] <ids...>   remove scenarios permanently
  help [command]              display help for command

For a regression suite, exclude is the safer way to retire a scenario: it stays on disk and out of runs, and include brings it back once its flakiness is fixed. rook scenarios list shows what would run, and what could not, before a run starts.

Rook CLI flags a scenario as flaky from its history across runs. Within a run, the scenario's repeat field sets how many samples it takes, and the Agent Assurance scenarios guide advises raising it only where several samples answer a real reliability or performance question.

Most eval and observability tools score what your agent said and recorded. Agent Assurance checks what the run changed, and reports what it could not verify. Eval tools rerun cases you wrote or synthesized from documents and traces, while Agent Assurance derives its scenarios from the agent's code; like eval tools, it uses model judges, grades what the agent says as well as what it did, and runs in CI.

Agent AssuranceAI eval toolsLLM observability
Test casesFrom your code or specWritten or synthesizedFrom production traces
Tool callsAgainst declared toolsAgainst your listsLogged, optionally scored
Side effectsFiles, artifacts, probesScripted per taskTrace data only
GradingClaimed actions aren't proofLLM judge or code checksLLM judge or human review
Adversarial testsGenerated by defaultAdd-on in some toolsNot generated
When it runsBefore release, in CICI and live trafficProduction, plus CI
Unverifiable resultsReported separatelyErrors or opt-in skipsLeft unscored

The eval and observability columns describe each category's default approach, not any single product.

Rook CLI installs from npm on Node.js 22 or newer, and the Rook CLI installation guide covers Homebrew and the shell installer, which bring their own Node runtime:

npm install -g @testmuai/rook
rook --version

To drive it from Claude Code, install the skill, then ask for the regression run in a /rook request:

npx @testmuai/rook-skill@latest install --agent claude-code
/rook Rerun this agent's current scenarios on the staging profile as a regression run. Before invoking the agent, show me the plan and which writes it could make in staging. After the run, compare it with the previous run: list newly failing, newly fixed and flaky scenarios, and any whose definition changed, with each failed criterion and its evidence.
Note

Note: Agent Assurance keeps scenarios, runs and evidence as plain, committable files in your project, so each release's baseline can sit next to the code it tested. Check the evidence for secrets and personal data before you commit it.

Conclusion

Start agent regression testing with the change your team ships most often. Pick the scenarios that exercise it hardest, run each several times on the current release to record a baseline, and rerun them before the next change ships.

Sort the results into newly failing, newly fixed, flaky, changed, still failing and unverified, and let newly failing scenarios, plus any still-failing one nobody has accepted, block the release. If your agent runs on an open-source harness, expect changes you did not make: Ben Sghaier and colleagues counted OpenCode at 18.0 releases per week.

Then move the same gate into your pipeline, following the guide to run Agent Assurance in CI/CD, with every run pointed at staging.

Author

...

Samyak Goyal

Blogs: 30

  • Linkedin

Samyak Goyal is a Senior Member of Technical Staff at TestMu AI engineering Kane CLI, the command-line tool that runs browser automation from the terminal, where a flow described in natural language executes in a real Chrome browser and returns pass or fail with shareable proof. He is a backend engineer with 4+ years of experience, previously an SDE at Innovaccer, where he built APIs, introduced Kafka, and cut deployment from weeks to hours. Samyak also builds multi-agent systems, skill-orchestration frameworks, and a personal copilot that indexes 200+ microservice repositories.

Reviewer

...

Srinivasan Sekar

Reviewer

  • Linkedin

Srinivasan Sekar is Director of Engineering at TestMu AI (formerly LambdaTest), where he leads engineering and open-source initiatives behind the Selenium and Appium automation grid and owns TestMu AI's MCP Server. A committer to Appium and a contributor to Selenium, WebdriverIO, Taiko, and AppiumTestDistribution, he brings over 15 years of experience in quality engineering and open-source technologies. He is the author of the Apress book 'The MCP Standard: A Developer's Guide to Building Universal AI Tools with the Model Context Protocol,' a Certified Kubernetes and Cloud Native Associate, and an international conference speaker. Before TestMu AI he spent over eight years at Thoughtworks as a Principal Consultant and Quality Architect. Srinivasan holds a B.Tech in Information Technology from Anna University.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Agent Regression Testing FAQs

Did you find this page helpful?

More Related Learning Hubs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests