Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Use Cases
- /
- AI Test Maintenance Without False Passes
AI Test Maintenance Without False Passes
Automated test repair fixes stale tests and can also hide real defects. How to scope re-authoring, review the draft it produces, and measure the gate.
Published on:
On This Page
A suite that needs a day of repair after every release stops being run. A suite that repairs itself silently stops being believed. Most teams adopting automated maintenance walk from the first problem straight into the second without noticing the trade.
This page is about the settings and the review habits that keep automated repair on the useful side of that line, using the self-maintenance strategies in KaneAI as the worked example.
TL;DR
Automated test repair keeps a suite running when the interface moves, and the same mechanism can rewrite a test around a genuine defect. Treat it as a proposal engine: scope what it may rewrite, keep approval on wherever a wrong repair is expensive, and measure the drafts a reviewer declines rather than the runs that went green.
What Is the Difference Between Healing a Locator and Re-Authoring a Test?
- Locator recovery: repairs the address of one element at runtime, changing nothing in the test and leaving nothing to approve. Low risk by construction, because it changes where a control is, not what the step asserts.
- Re-authoring: rewrites the objective that failed and every objective after it, then saves the result as a draft version. This is a content change, so the new steps and new assertions are generated rather than authored by a person.
- Blast radius: a failure in the third of five objectives leaves the first two as they replayed and rewrites three, so both the credit cost and the review surface grow with how much of the test sits after the failing step.
- The D3 to D4 line: repair a changed control, a moved element or a new step in the journey; refuse when the option the test selects has disappeared. Above that line the application moved, below it the application broke.
Does Automated Test Repair Actually Work?
Well enough to be useful, not well enough to be unsupervised. An industrial study of autonomous test repair reported a 70% convergence rate that falls to 50% once the runs that converged by weakening an assertion or deleting a test case are excluded. Twenty points of the headline number were repairs that passed by testing less.
Can a Self-Healing Test Hide a Real Bug?
It can, because a step fails the same way whether a button moved or the feature broke, and re-authoring answers both with a green run. The approval gate is the one checkpoint that tells them apart, which is why a re-authored version is held as a draft by default rather than becoming current.
Why Do Recorded Tests Go Stale?
A recorded test replays the steps it was authored with. The application does not hold still, and each kind of change breaks the replay in a different way and deserves a different repair.
| What changed | How the replay fails | The repair that fits |
|---|---|---|
| An element moved or was renamed | The locator no longer matches, though the step is still the right step | Locator recovery at runtime, with no change to the test |
| A flow gained or lost a step | The sequence no longer matches the product, so later steps land on the wrong screen | Re-authoring the objective, because the recorded sequence is now wrong |
| The page is genuinely dynamic | No recorded sequence survives long enough to be worth maintaining | Authoring from the objective on every run instead of replaying |
| Timing or environment wobbled | The same test passes on a second attempt with nothing changed | A retry, because there is nothing in the test to fix |
| The feature broke | The step fails because the expected behaviour is no longer there | None. This failure is the test working correctly |
The last row is the one automation cannot see from the inside. Every row above it looks identical to the row below at the moment of failure, which is the whole problem this page is about.
When Is a Passing Test Not a Working Test?
A repair is judged by whether the test passes afterwards. That is a weak oracle, and the research on automated repair has been saying so for a decade.
The most direct evidence comes from an industrial case study of an autonomous multi-agent testing system. Across 300 consecutive execution reports covering 636 test-case executions, Practical Limits of Autonomous Test Repair reports a 70% repair convergence rate at the scenario-family level, and a semantically strict rate of 50% once the families that converged by weakening an assertion or deleting a test case are excluded.

Twenty points of that headline number were repairs that made the test pass by making it test less. The study names the two mechanisms it observed.
- Assertion weakening - one family converged by relaxing a strict equality check into a truthiness check, which passes for almost any value the application returns.
- Silent test-case deletion - another converged because the failing case stopped existing, which removes the failure and the coverage together.
- Convergence took several attempts - only 10% of scenario families succeeded on the first attempt, at a mean of 4.4 repair iterations, so most repairs are the product of a loop rather than a single confident edit.
The same weakness shows up in program repair, where the test suite is the thing being satisfied. Work on identifying patch correctness in test-based program repair starts from the observation that test suites in practice are often too weak to guarantee correctness, and that existing approaches generate a large number of incorrect patches. Evaluated on 139 patches from five repair systems, its heuristic blocked 56.3% of the incorrect ones without blocking any correct patch.
| Study | What it measured | The finding that matters here |
|---|---|---|
| Autonomous test repair, 2026 | 300 execution reports, 636 test-case executions, 10 scenario families | 70% converged, 50% converged without weakening an assertion or deleting a case |
| Patch correctness in program repair | 139 patches from five automated repair systems | Test suites are often too weak to guarantee correctness; a heuristic blocked 56.3% of incorrect patches without blocking a correct one |
Read across the two, the useful conclusion for a QA team is narrow and firm. Automated repair earns its place as a proposal engine, and something outside the loop has to decide whether the proposal is a fix or a cover-up.
What Exactly Gets Repaired, the Locator or the Test?
Vendors use "self-healing" for two mechanisms with very different risk profiles. Separating them is the first governance decision, because only one of them changes what your test asserts.

- Locator recovery is low risk by construction - it changes the address of an element, not the intent of the step. KaneAI's Auto-Heal tries the other locators captured for that element, then rebuilds one from the step's original natural language instruction, and creates no version.
- Re-authoring is a content change - it rewrites the objective, so the new steps and the new assertions are generated rather than authored by a person. That is why it produces a version and why the version waits.
- They can both act on one run - locator recovery runs first, and re-authoring is what handles the step it could not recover.
If a tool cannot tell you which of these two it is doing, treat everything it does as the second one. The reviewing effort you save by assuming the first is exactly the effort that catches a weakened assertion.
Two categories is the minimum. The useful version of the question is how much of the original test a system is allowed to reinterpret, and that runs to six levels.
| Level | What failed | What the system is permitted to do | The risk it introduces |
|---|---|---|---|
| Interaction recovery | The element exists but is not actionable yet | Bounded wait, scroll, re-query, retry the same action | Masking a real responsiveness defect as a timing wobble |
| Locator healing | Same element, different address | Substitute a locator from the fallback set | Selecting a similar but wrong control |
| Semantic element healing | No stored locator resolves at all | Rebuild the target from the step's original instruction and page context | Ambiguity between two plausible targets |
| Step healing | The action now needs a different sequence | Insert, reorder or change UI actions | Changing what the test does, not just how it does it |
| Objective healing | A whole part of the journey is stale | Re-author the objective and everything after it | A false green across a section nobody reviewed |
| Dynamic execution | Nothing. The recorded sequence is treated as disposable | Plan the journey from intent on every run | Non-determinism, and failures that are hard to reproduce |
Approval requirements should climb with that table rather than sit flat across it. The cheap deterministic levels earn their autonomy; the interpretive ones have not.
Order matters as much as permission. Running the expensive interpretive level first is how teams end up paying an agent to solve a problem a fallback locator would have solved, and research on locator repair has been arguing the same for years: Brisset, Rouvoy, Seinturier and Pawlak built Erratum on tree matching precisely because scanning every element on a changed page gets slower and less accurate as interfaces grow.
How Much of a Test Does One Failure Rewrite?
Re-authoring is not a surgical edit to the step that broke. The unit is the objective, and the scope runs to the end of the test.

That shape has three consequences worth planning around, and they all point the same way: a failure early in a long test is the expensive one.
| Consequence | Why it follows | What to do about it |
|---|---|---|
| Review surface grows | Objectives that never failed are rewritten because they sit after the one that did | Keep objectives small and ordered so a failure rewrites less of the test |
| Cost grows with position | Re-authoring consumes authoring credits, and a trigger covers everything after the failure | Put the steps most likely to break late in the test where that is possible |
| Earlier results survive | Objectives that already ran keep the result they replayed | Trust the pre-failure portion of the run, and review only from the failure onward |
A run in which nothing fails to replay does not trigger re-authoring at all, so the cost is a function of instability rather than a standing charge. The one exception is authoring from objectives on every run, which pays that cost whether or not anything broke.
Which UI Changes Should a Test Absorb?
Most healing policies are written in terms of confidence, which describes the system's certainty and says nothing about whether the change was one a test should follow. Classifying the drift instead gives you a rule that survives contact with a real interface.

The boundary between D3 and D4 is the one worth writing into policy. Above it the application moved and the test should follow; below it the application broke and following it produces the false pass.
- D0 and D1 are free - a renamed class or a moved node changes nothing a user could observe, and a repair here needs no ceremony beyond a log line.
- D2 is where agents start earning their cost - when a dropdown becomes a set of cards there is no old locator left to repair, so pattern matching has nothing to work with and only intent does.
- D3 needs a named reason - a new step in the journey is either a product decision somebody made or a bug, and a reviewer should be able to point at the release that caused it.
- D4 and D5 are findings - the correct output is a defect report and a red run, and any system that repairs its way past them is working against you.
A healer that reports only pass or fail cannot express this distinction, which is the practical argument for carrying the drift class into the run result rather than leaving it in a log.
What Should a Repair Never Be Allowed to Change?
The D3 to D4 boundary is only enforceable if the test states what it is proving separately from how it proves it. Splitting the two turns a cultural norm into something a system can check.

The practical form of the immutable half is a short list of user-visible facts that must still hold after any repair. For a shopping journey it is close to trivial to write down.
- The product identity - the item in the basket is the item the objective named, not the one nearest the pointer.
- The chosen variant - size M stays size M. A repair that settles for L because L was the easiest chip to find has changed the test.
- The quantity - one remains one, which catches the class of repair that clicks twice to make a stubborn control respond.
- The visible confirmation - the basket badge and the line item still appear, so success is measured on the screen rather than on a method returning without an exception.
Checked after the repair rather than before it, that list is the difference between a test that adapted and a test that gave up. It also gives a reviewer something concrete to disagree with, which a confidence score never does.
Confidence and risk belong on separate axes for the same reason. Being 99% certain that a control is the delete button is not an argument for clicking it, and a policy that folds one number into the other loses exactly the distinction that matters on destructive actions.
| Certainty the target is right | Cost of getting it wrong | What the system should do |
|---|---|---|
| High | Low, such as navigation or a read-only view | Repair and continue, recording what it did |
| High | Medium, such as a state change the user can undo | Repair, then replay from a clean session to confirm |
| High | High, such as deletion, payment or a permission change | Stop and ask. Certainty is not authority |
| Medium | Any | Propose the repair, do not perform it |
| Low | Any | Fail without acting, and say which candidates were considered |
The third row is the one that separates a governed system from an eager one. It is also the row most healing configurations do not have, because a single threshold cannot express it.
Note: Adaptive Heal, Dynamic Test and Retry on Failure sit under one Self-maintenance setting in KaneAI, with approval off by default so a re-authored version waits in Version History. Start free and set a strategy on one run before you set a project default.
What Does One Repair Actually Look Like?
Everything above is policy. This is the same policy applied to one checkout test across two releases, where the interface change is identical and the correct verdict is opposite.
The test has five objectives. The third selects a delivery option.
objective_3: Choose express delivery
recorded_steps:
- click: "#delivery" # a <select> element
- select: "express"
- assert: text "Express delivery" is visible
invariants: # what this objective proves
- selected_delivery == "Express"
- delivery_cost is visible
- basket_total updatesRelease A: the checkout is redesigned
A designer replaces the dropdown with three cards. No selector in the recorded steps resolves, the step fails, and re-authoring produces this draft.
objective_3: Choose express delivery
- click: "#delivery"
- select: "express"
+ click: card labelled "Express delivery"
assert: text "Express delivery" is visibleRun the four reviewer checks against it and every one passes: the assertion is unchanged, nothing was dropped, the failure has a matching design ticket, and the rewrite is confined to the control that moved. The invariants hold after the repair, because Express is still what ends up selected. Approve.
Release B: the same failure, a different cause
A pricing change removes Express from three postcodes, and a bug removes it from all of them. The step fails exactly as before, and the draft comes back like this.
objective_3: Choose express delivery
- click: card labelled "Express delivery"
+ click: card labelled "Standard delivery"
- assert: text "Express delivery" is visible
+ assert: a delivery option is selectedThe run goes green. The suite reports no regression. Two of the four checks fail: the assertion was weakened from a specific string to an existence check, and no ticket explains why Express disappeared. The invariant selected_delivery == "Express" is false. Decline, and raise the bug.
| Signal | Release A | Release B |
|---|---|---|
| What the run reported | Passed after a repair | Passed after a repair |
| Drift class | D2, representation | D4, meaning |
| Assertion | Unchanged | Weakened to an existence check |
| Invariants after repair | All hold | selected_delivery fails |
| Correct verdict | Approve | Decline and raise a defect |
Nothing in the run distinguishes these two. Same failing step, same green result, same draft waiting in Version History. The only thing separating a legitimate repair from a shipped bug is a person reading the second diff against a list of facts written before either release.
This is also the argument for writing the invariants down first. A reviewer looking at Release B without them sees a plausible rewrite of a test they did not author, under time pressure, in a queue of other drafts.
Which Strategy Fits Which Suite?
The choice is usually made once, at project level, for suites that have nothing in common. Matching the strategy to what the suite is for produces a better answer than a single organization-wide default.

| If the suite is | Choose | And set approval to | Because |
|---|---|---|---|
| The release gate | Off | Not applicable | A gate exists to stop the release. A gate that repairs itself past a failure is not a gate |
| Regression on a UI that churns | Adaptive Heal | Off, so drafts queue for review | The tests still describe the right thing, and the recorded steps are what went stale |
| Built on pages that change constantly | Dynamic Test | Off on anything release-relevant | A recorded script is a liability when no sequence survives to the next run |
| Failing on timing and environment | Retry on Failure | Not applicable, nothing changes | There is nothing in the test to repair, so re-authoring would rewrite a correct test |
| Compliance or audit evidence | Off | Not applicable | The value of the record depends on a person having authored what it asserts |
Retry on Failure deserves one caution that is easy to miss. Retrying is the right answer only when the failure is genuinely non-deterministic, and using it to paper over a real defect produces the same false pass by a slower route, which the flaky test guide covers in more depth.
What Should a Reviewer Actually Check?
"Review the draft" is not a process until someone writes down what would make them decline one. These five checks come straight from the failure modes the research documented.
- Does it still assert the same thing? Compare the draft against the current version and read the assertions first, not the steps. An equality check that became a presence check is the documented failure mode.
- Did anything get dropped rather than fixed? A shorter objective that passes is a coverage loss wearing the costume of a repair.
- Is there a defect report for the original failure? If the step failed because the product broke, the correct outcome is a declined draft and a raised bug, not a new version.
- Does the rewritten portion match a real product change? Someone should be able to name the release or ticket that made the old steps wrong.
- Is the rest of the rewrite incidental? Objectives after the failure are rewritten because of their position, so changes in them need the same reading even though nothing in them failed.
KaneAI supports this with the mechanics rather than leaving it to discipline.
- The draft is attributed to the strategy - not to the person who started the run, so nobody inherits authorship of a change they did not make.
- One test case produces one draft - a run covering the same case under several browser versions or operating systems still yields a single draft to read, because it is one test case.
- A waiting draft cannot be edited - it has to be approved or declined first, which stops a half-reviewed rewrite from being patched into acceptability without a verdict.
- Approving regenerates the exported code - so the framework code your pipeline runs matches the version that was approved rather than the one before it.
- Declining costs only the run you already paid for - it discards the draft and leaves the current version untouched, which makes declining the cheap option rather than the disruptive one.
How Does a Good Repair Become a Bad Baseline?
A repair that worked once worked under the exact conditions that produced the failure. Promoting it on that evidence is how a lucky recovery becomes the reference point every later run is measured against.
The failure mode compounds in any system that learns from successful runs, and none of the steps look wrong on their own.
- A repair picks the wrong but plausible target.
- The step completes, so the run is recorded as successful.
- That successful run becomes the new reference for what normal looks like.
- Later repairs measure themselves against the wrong target and agree with it.
- The original defect is now invisible, and the suite is confidently wrong.
The fix is not to stop learning from successful runs, which would throw away the thing that makes healing cheap. It is to make promotion conditional on the user-visible facts rather than on the absence of an exception.
| Practice | What it does | What it prevents |
|---|---|---|
| Replay in a fresh session | Run the repaired journey from the start in a clean session before accepting it | A repair that only works from the mid-journey state the failure left behind |
| Check the invariants, not the exit code | Confirm the user-visible facts after the repaired step, not that the click returned | The wrong-element repair that technically succeeded |
| Gate promotion on that check | Only a run that satisfied its invariants may become the new reference | A poisoned baseline that legitimises every later repair |
| Remember rejected targets | Record what a reviewer declined and penalise it in later matching | Oscillation, where the system re-proposes the target you rejected last week |
The same logic is why a run that passed after a repair should not be filed under the same result as a run that passed outright. Collapsing the two hides the number a team most needs to watch: how much of the suite is now standing on repairs nobody read.
Where Should the Strategy Be Decided?
A maintenance strategy set once at the top of an organization is a policy decision disguised as a toggle. The inheritance rules decide who actually owns it.

- Set the organization level to the safest option - it is a default for projects that have not thought about it yet, and the safest default is the one that fails loudly.
- Expect projects to diverge and let them - a project follows the organization until someone changes it there, after which it keeps its own value and later organization changes no longer overwrite it.
- Treat run-level choices as experiments - a run starts from the project value and can override it, and changing it in a run never writes back to the project or the organization.
- Remember it is not retroactive - turning self-maintenance on applies from that point forward, and runs that already exist are not eligible whatever the setting now says.
One practical habit prevents most surprises. Confirm the strategy in Advanced Configurations before executing, because that is where the organization, project and run-level values resolve, and the test run panel shows the resolved strategy alongside the instance count and concurrency.
How Do You Know the Gate Is Working?
The obvious metric is the wrong one. A rising pass rate after enabling automated repair is exactly what you would see if the repairs were correct, and also exactly what you would see if they were hiding defects.
| Measure | How to compute it | What it tells you |
|---|---|---|
| Draft approval rate | Drafts approved divided by drafts produced | How often the repair was judged correct. A rate at or near 100% is a review problem, not a quality result |
| Declines that became defects | Declined drafts with a linked defect divided by declined drafts | The gate's actual yield. Each one is a bug that a green run would have buried |
| Time to verdict | Median hours a draft waits in Version History | Whether review is a step in the workflow or a queue nobody owns |
| Trigger rate | Runs that triggered re-authoring divided by eligible runs | How unstable the suite really is, and what the credit spend is tracking |
| Objectives rewritten per trigger | Objectives re-authored divided by triggers | The real blast radius in your suite, which is a function of where failures land |
| Escapes traced to an approved heal | Production defects whose coverage was re-authored before release | The guardrail. One of these is a signal to tighten approval, not to review harder |
| Eligible share of the suite | Eligible test cases divided by total test cases | How much of the maintenance problem this strategy can address at all |
The second row is the one to put in front of a stakeholder. It converts a governance step that looks like overhead into a count of defects that would otherwise have shipped, which is the argument for keeping approval switched on.
One derived measure is worth more than the rest of the table combined. Call it the false-green rate: repaired passes that no longer prove what the test was written to prove, divided by all repaired passes.
It reframes what a good result looks like. A system recovering 98% of broken tests with a 5% false-green rate is less trustworthy than one recovering 90% with almost none, because the first is retiring coverage while reporting success.
Give the pipeline more than two words for what happened
- Passed outright - the recorded steps replayed and nothing was repaired. This is the only result that needs no further thought.
- Passed after a retry - nothing in the test changed, but the count belongs in the result, because a step that needs three attempts is reporting something about the product.
- Passed after a locator repair - safe to ship, worth counting, and a rising trend says the interface is churning faster than the suite is being maintained.
- Passed after a re-plan - the journey changed shape. This is the state that should never be collapsed into a plain pass.
- Held for review - the repair was plausible but crossed a risk threshold, so the run has an answer and a person still owes it a verdict.
- Failed on meaning - a repair was possible and was refused because it would have changed what the test proves. This is the system working correctly.
The last state is the one most tools cannot express, and it is the one that makes a healing system trustworthy. A refusal to repair is a finding, and a pipeline that can only say pass or fail has nowhere to put it.
Where to set the first thresholds
None of these are benchmarks, because no published dataset covers them. They are starting lines, chosen so that crossing one prompts a conversation rather than an alarm, and meant to be replaced by your own numbers after a quarter.
| Measure | Starting line | What crossing it should trigger |
|---|---|---|
| Draft approval rate | Above 95% | Read ten approved drafts at random. Near-total approval usually means the queue is being cleared, not read |
| Declines that became defects | Zero over a quarter | Either the product is unusually stable or the gate is ceremonial. The mutation benchmark settles which |
| Time to verdict | Beyond one working day, median | Give the queue an owner. Drafts that age get approved in batches, which is where weakened assertions enter |
| Trigger rate | Above 20% of eligible runs | Fix the tests rather than the repairs. At that rate you are paying credits for instability you could author out |
| Objectives rewritten per trigger | Above three | Split the long tests. Failures are landing early and rewriting most of the journey behind them |
| Escapes traced to an approved repair | One, ever | Turn auto-approve off everywhere it is on, and re-run the benchmark before turning it back on |
The second row is deliberately counterintuitive. A gate that has never declined anything is not evidence of a healthy suite, and treating zero as a good score is how a review step becomes a rubber stamp nobody notices.
What Changes by Team and Domain?
The policy is the same everywhere. What moves is how expensive a wrong repair is, and that should decide the settings rather than the size of the team.
| Context | What a wrong repair costs | Where to set the line | The invariant nobody should skip |
|---|---|---|---|
| Payments and checkout | A charge of the wrong amount, to the wrong instrument | Repair below D3, refuse above it, approval always manual | Amount, currency and instrument, asserted from the screen after the repair |
| Regulated and audited systems | A test record that no person authored, in an evidence pack | Repair off entirely on anything that produces evidence | Who authored the current version, and when a human last approved it |
| Permissions and access control | A test that proves the wrong role can do the thing | Refuse any repair that changes which account or role is used | The acting identity, checked after the repair rather than assumed from setup |
| High-churn consumer UI | Little, if coverage survives | The best fit for automatic repair up to D3, with sampled review | The user-visible outcome of the journey, not the individual clicks |
| Internal tools and admin panels | Almost nothing, until a destructive action is involved | Automatic repair, with healing disabled around delete and bulk actions | That the destructive step targeted the record the test named |
| A team of one or two | Review capacity, which is the scarce resource | Narrow eligibility rather than loosening approval | Fewer tests under repair, each one actually read |
The last row is the one small teams get wrong most often. Facing a review queue they cannot service, the instinct is to turn auto-approve on, which removes the gate precisely because there was no capacity to staff it.
Narrowing what is eligible is the better trade. Ten tests under repair with every draft read beats two hundred under repair with none, because the second arrangement produces confidence without evidence.
How Does an Agentic Repair Loop Actually Work?
Once repair moves past locator substitution it stops being a lookup and becomes a loop: try something, look at what happened, try again. What separates the implementations is not the reasoning, it is where the loop is made to stop.
Playwright's test agents document the shape openly. Its healer replays the failing steps, inspects the current UI for equivalent elements or alternative flows, suggests a patch such as a locator update, a wait adjustment or a data fix, and re-runs until the test passes or guardrails stop the loop.
One detail in that description is worth more than the rest. If the healer concludes the underlying functionality is genuinely broken, it skips the test rather than continuing to repair, which is the same refusal the drift envelope asks for, expressed as a product behaviour.
| Loop property | Question to ask of any agentic healer | Why it decides whether you can trust the result |
|---|---|---|
| Unit of repair | Does it patch a locator, a step, a whole objective, or the test source? | It sets how much can change while the result still reads as one pass |
| Stop condition | What ends the loop besides the test going green? | A loop that only stops on success will eventually succeed at the wrong thing |
| Iteration budget | How many attempts, and is the count reported? | A repair that took nine attempts is telling you something a green tick hides |
| Refusal behaviour | What does it do when it decides the product is broken? | Skipping, failing and silently passing are three very different answers |
| Persistence | Does the repair end with the run, or become the test? | Runtime recovery and a rewritten test need different approval bars |
| Determinism | Does the same failure produce the same repair twice? | A non-reproducible repair cannot be reviewed, only accepted or rejected |
The refusal row is where most evaluations stop short. A healer that skips a test it cannot honestly fix has produced a finding, and a suite that carries skipped tests without surfacing them is not much better than one carrying false passes.
What the documented approaches actually repair
| Approach | Unit repaired | Leaves behind | Where it stops |
|---|---|---|---|
| Locator recovery | One element's address | Nothing. No version, no approval | When no stored locator and no rebuild from the instruction resolves |
| Objective re-authoring | The failing objective and every one after it | A draft version awaiting a verdict | At the approval gate, by design |
| Authoring on every run | Every objective, whether or not anything failed | A draft version, every run | Nowhere automatically, which is why it needs the tightest scoping |
| Whole-test retry | Nothing. It runs the same test again | Nothing to approve | At the retry cap, which is the only protection it has |
| Source-patching agents | The test file: locators, waits, data | A patch, and sometimes a skipped test | At guardrails, or when it judges the feature genuinely broken |
| Baseline-matching mobile healers | A mobile locator, against a snapshot from the last passed run | A log of original and healed selectors, or a suggestion | At low confidence, where it suggests rather than heals |
| Tree-matching research | A broken locator, by structural comparison | A relocated element, deterministically | At the locator. It never reasons about the journey |
Read down the third column rather than the first. Approaches that leave nothing behind cannot produce a false pass that outlives the run; approaches that leave a version or a patch can, and those are the ones that need the gate.
Several vendors publish less than this about where their repair stops, which is itself an evaluation signal. A tool that will not tell you its unit of repair or its refusal behaviour cannot be governed with the policy above, whatever its recovery rate looks like in a demo.
Both belong in the result rather than in a log, which is the same argument as testing agentic AI beyond pass fail. For how agents fit into UI automation more broadly, agentic testing in UI automation covers the surrounding workflow.
Does Any of This Change on Mobile?
The governance holds. What changes is how a target is identified and how many ways a step can fail without the interface having drifted at all.
TestMu AI's Smart Heal shows the shape of a mobile healer. It needs at least one passed run to capture a baseline snapshot of every locator in the script, compares element attributes, hierarchy and visual cues when a later locator fails, retries the step with the healed locator, and logs the original alongside the recovered one. Where it cannot identify an alternative confidently, it records a suggestion instead of forcing a heal.
One documented behaviour deserves the attention of anyone running it: the baseline updates automatically after each successful run. That is efficient, and it is the exact mechanism the poisoning section describes, which makes the invariant check before promotion a practical safeguard rather than a theoretical one.
Element not found is not one failure
- Not rendered yet - the screen is still settling, which is a bounded wait rather than a locator problem.
- Off-screen - the control exists below the fold, so the answer is a controlled scroll.
- In another window - elements belonging to a keyboard or overlay may sit outside the default hierarchy entirely.
- The hierarchy moved - the genuine locator-healing case, and the only one of these five that a healer should repair.
- The control is gone - a finding, and the case where healing is the wrong instinct.
Sending all five to an AI matcher produces repairs for problems that were never about identity. Localisation makes the point sharpest: Appium's XCUITest locator strategies warn that an element's name attribute "is commonly rendered in the application GUI, which means it is likely to change depending on the application language".
A button whose label moves from Continue to Continuer has not disappeared. A healer that treats the visible label as identity will either fail a correct build or heal onto whatever else is nearby, and both outcomes are worse than the locale switch that caused them.
How Would You Prove a Healer Is Safe?
Vendor demos ask whether the test recovered. The question that decides whether you can trust a healer is whether it recovered the right behaviour, and that needs a test suite of its own: deliberate changes to your interface, each with a known correct outcome.
Two families, and the second is the one everybody skips.
| Change you introduce | Family | The healer passes this test only if it |
|---|---|---|
| Rename an id or class | Structural | Repairs it without comment |
| Wrap the control in extra markup, or move it on the page | Structural | Repairs it without comment |
| Replace a dropdown with an equivalent set of cards | Structural | Re-plans, then confirms the same option ended up selected |
| Delay the control so it appears late | Structural | Waits, rather than rewriting a locator that was never wrong |
| Remove the option the test selects, leaving a near neighbour | Semantic | Refuses, and reports the option as missing |
| Delete the expected confirmation message | Semantic | Fails, rather than relaxing the assertion that looked for it |
| Swap the meaning of two similar controls | Semantic | Declines to choose between them |
| Let every click succeed but land the journey somewhere else | Semantic | Fails on the outcome, not on any individual step |
A benchmark built only from the first family makes an over-eager healer look excellent, because every mutation in it rewards repair. The second family is where a healer earns trust, and refusing correctly should score as a pass rather than a miss.
Eight mutations against one representative journey is enough to rank tools, settle an internal argument about auto-approve, or give a team the evidence to raise a threshold it had been guessing at.
Where Does This Approach Stop?
Automated re-authoring covers a narrower slice of a suite than the phrase "self-healing tests" suggests, and the boundaries are worth knowing before a rollout plan depends on them.
- Eligibility is narrow - in KaneAI, re-authoring applies only to test cases using New Experience with a Chrome browser configuration, and everything else replays its recorded steps whatever is set.
- One strategy at a time - selecting one clears the others, so a run cannot both re-author and retry, and a suite needing both behaviours needs two runs.
- It does not reduce the review workload to zero - it converts a rewrite into a review, which is a significant reduction in maintenance rather than the removal of it.
- It repairs tests, not test design - an objective that was ambiguous when a person wrote it will be re-authored into something equally ambiguous.
- A stale repository is still stale - repair keeps individual tests running and does nothing about the duplicates and dead cases around them, which is the separate problem covered in duplicate test case detection with AI.
The honest framing for a rollout is that this replaces one kind of work with a smaller kind of work. Teams that budget for the review time get the benefit, and teams that assume the review disappears are the ones who find the weakened assertions months later.
The ways a healing programme goes wrong, and what stops each one
| Failure mode | Why it happens | What prevents it |
|---|---|---|
| Wrong-element repair | Two controls look and read alike | Checking the user-visible outcome after the action, not that the action returned |
| Retry masking | A real responsiveness defect is treated as a timing wobble | A hard retry budget, with the attempt count carried into the result |
| Assertion weakening | Relaxing the check is an easier route to green than fixing the journey | A higher approval bar for assertion changes than for locator changes |
| Repair oscillation | Two candidates are equally plausible, so the system alternates | Recording rejected targets and penalising them in later matching |
| Responsive ambiguity | The same action exists twice, once per viewport | Carrying viewport and device into the target description |
| Fragile repairs | The system optimises for resolving now, not for surviving the next release | Scoring candidates on durability, so repairs do not rebuild the debt |
Each row has the same shape: something cheaper than the correct answer was available, and the system took it. That is what a budget and a review gate exist to interrupt.
Conclusion
Pick the one suite that costs the most to maintain, set Adaptive Heal on a single run with approval left off, and read the first draft it produces against the current version before approving it. That single review tells you more about whether automated maintenance suits your application than any pass-rate chart will.
From there, keep approval on wherever a wrong repair is expensive, and start counting declined drafts that turned into defects. The Adaptive Heal and Dynamic Test documentation covers the settings, and the feature walkthrough covers the screens they live on.
Author
Abhishek Mishra is a Technical Product Manager at TestMu AI (formerly LambdaTest), where he owns Test Manager, the test management product. He has over 8 years of experience in product management and market analysis, spanning AI-native software testing, product strategy, and analytics. On TestMu AI, he authored guides on test management and test case management. Previously, he served as the Product Lead at IndiaClan and co-founded Gartley618 Technologies, a firm focused on quantitative trading and blockchain. He holds a B.Tech degree.
Reviewer
Shantanu Wali is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he owns several product lines across the testing platform, including the Real Device Cloud and the Digital Experience Testing Cloud. He has also contributed significantly to the development and scaling of KaneAI, TestMu AI's flagship GenAI-native testing agent that uses natural language to make software testing faster and more reliable in this AI era. He brings 7+ years of experience across software development and product management, starting as a backend developer at Infosys building solutions for Fortune 500 clients. Shantanu holds an MBA from IIM Calcutta and a B.Tech in Mechanical Engineering.
AI Test Maintenance FAQs
Did you find this page helpful?
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests





