Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Two-Way Doors for QA: Keep Every Release Reversible
Two-Way Doors for QA: Keep Every Release Reversible
Two-way doors keep every release reversible. How QA teams test feature flags, dark launches, and rollbacks, building on Google's Two-Way Doors TotT episode.
Published on:
Teams that act quickly do not shun risky decisions; rather, they ensure that such decisions can be reversed and that their tests prove such a reversal is safe.
The article on the Google Testing Blog today, "Two-Way Doors: Don't Code Yourself into a Corner" by Nimit Khandelwal and Chris Kennelly, makes a straightforward point. In large codebases, small decisions can lead to hidden dependencies, and if no provision is made for an exit, those decisions gradually become irreversible. The solution suggested by the authors is to deliberately include escape hatches such as feature flags, narrow APIs, separate data formats, and dark launches.
The post comes from Google's long-running Tech on the Toilet series and is adapted from Performance Tip of the Week #87. It's written for engineers designing systems, but it is just as much a testing story. An escape hatch only works if you can tell, quickly and with confidence, when to use it.
One-Way vs Two-Way Doors, for QA Teams
Google classifies decisions into two categories. A one-way door is uncommon, high-consequence, and difficult to reverse, for example, when selecting a database backend. A two-way door is easy to reverse, such as when naming a local variable. The problem is that two-way doors can gradually become one-way doors once some code has been written around them.
For QA teams, what makes a door two-way is evidence, not just architecture:
| Escape hatch (from Google's post) | What it protects | What testing has to show |
|---|---|---|
| Feature flags | Releasing separately from deploying | Both flag states work, and flipping back is safe |
| Limiting API exposure | Room to refactor interfaces | Contract tests catch breaking changes before anyone widens access |
| Decoupling data formats | Cheap in-memory changes | Old and new data still read correctly, both ways |
| Dark launches | Validating at production scale | New logic matches old output under real traffic |
If you can't verify a rollback, you don't really have a two-way door. A flag nobody has tested in the "off" state is a one-way door that hasn't been found yet.
Feature Flags: Every Flag Doubles Your Test Paths
Feature flags are the most frequently used way to get out of a situation and are also the ones most often not tested; since each flag divides behavior into at least two paths, three separate flags on a single screen already result in eight combinations.
Most suites only test whichever state is currently on in staging. Three gaps follow from that:
- The "off" path rots. Nobody runs it until the rollback that needs it.
- Combinations break. Two flags that each work alone can conflict when both are on.
- Cleanup is risky. Deleting a stale flag is also a code change and requires the same level of coverage as shipping one.
The fix is to treat flag state as test input. Write a scenario once, run it across the flag states that matter, and keep the "off" path tested until the flag is removed.
Dark Launches: Check Before You Commit
In the case of a dark launch, the new logic runs alongside the old one while actual traffic is processed; its results are discarded, and only the telemetry is retained. Google refers to this as a method of verifying the correctness and performance at production scale before relying on the new path.
Thus, testing in a production environment becomes a purposeful element of the release process rather than a risk. A dark launch performs most effectively when three stages of checking surround it:
- Before deployment - end-to-end tests include both the previous and the new routes in the staging environment, across the browsers and devices your users actually use.
- During the dark launch - the telemetry system compares the old and new outputs, and scheduled end-to-end tests are run in either the canary or production environment to ensure that the user-facing process has not changed.
- Before you commit - the same suite will operate using the new path. A pass serves as your proof that you can go through the door, while a fail means you revert with nothing lost.
Testing doesn't stop at merging. It runs continuously until the change has earned its place, as Google's post implies without stating it outright.
Where KaneAI Fits
KaneAI turns verification into something you can run at every door, not only at merge.
| Two-way-door need | How KaneAI helps |
|---|---|
| Test both flag states without writing two suites | Write the scenario once in plain English and run it across environments, including canary, blue/green, and A/B setups (AI QA Agent) |
| Keep checking during a dark launch | Scheduled runs, on demand or at fixed intervals, against any environment (KaneAI) |
| Real-user coverage before you commit | Tests run on 3,000+ browsers and 10,000+ real devices through HyperExecute |
| Gate the pipeline on evidence | Kane CLI runs headless in CI and returns exit codes your pipeline can gate on |
| Keep the off path from rotting | Self-healing updates locators when the UI shifts and shows the diff for review |
| Check each change at the PR | Comment on a GitHub PR and KaneAI generates and runs tests from the diff (beta) |
| Verify AI features before rollout | Agent Testing for chat, voice and video agents |
KaneAI also has versioning with rollback for the tests themselves. Your suite stays a two-way door too.
A one-line start with Kane CLI:
npm install -g @testmuai/kane-cli
kane-cli run "verify checkout flow on staging with new-pricing flag off"A Two-Way-Door Testing Checklist
Use this before any flagged or dark-launched change ships:
- Every new flag has end-to-end tests for both the on and off states
- Flag combinations that share a screen or flow are tested together
- The rollback path (flag off, old format, previous API) passed within the last 24 hours
- Contract tests protect any API before its access is widened
- Old and new data formats read correctly, forward and back
- Dark-launch telemetry compares old and new output, with a stated threshold for divergence
- Your scheduled end-to-end tests run against canary or production while the rollout is happening
- Tests pass on the real browsers and devices your users have, not just Chrome on one laptop
- A failing test blocks the release automatically, so nobody has to remember to check
- When you remove an old flag, you test it as thoroughly as when you added it
Sign off on the decision only when every box is ticked.
Further Reading
From Google
- Two-Way Doors: Don't Code Yourself into a Corner, Google Testing Blog, Oct 5, 2026
- Performance Tip of the Week #87, the source of the Two-Way Doors episode
- Tech on the Toilet: Driving Software Excellence, One Bathroom Break at a Time, Google Testing Blog, Dec 2024
- Change-Detector Tests Considered Harmful, on tests that break on every change and make refactoring a one-way door
- Tests Too DRY? Make Them DAMP!, on readable tests that are easy to change
From TestMu AI
- KaneAI, a GenAI-native testing agent
- Agentic Testing with KaneAI
- Kane CLI and its docs
- HyperExecute
- Agent Testing and Agent Assurance
Author
Anmol Gupta is Vice President of Product Management at TestMu AI (formerly LambdaTest), driving HyperExecute, the test orchestration cloud that runs and accelerates automated test execution. He led the development of the Unified Test Execution Cloud Platform and now leads a 30-member cross-functional product organization across product lines contributing $7M+ in revenue. He brings over nine years of experience and previously co-founded the SaaS company Timble as CTO, where he grew the team from 5 to 40 and launched an AI KYC platform that processed 600K+ applications in five months while cutting verification time from 12 minutes to under 30 seconds. Anmol holds an MTech and BTech from IIT Delhi.
Reviewer
Mayank Bhola is Co-Founder and Head of Products at TestMu AI (formerly LambdaTest), where he leads the entire product portfolio across KaneAI, Kane CLI, HyperExecute, SmartUI, the Real Device Cloud, Accessibility, and other software testing product lines. As an early Lead Architect he designed and built the company's flagship Tunnel technology from scratch, created the React-based automation platform, and architected the data-intensive pipelines and FAAS services that scale it. He brings more than 10 years of experience in software development and product engineering, with earlier roles as Head of Technology at Juggernaut Books and Senior Software Engineer at PressPlay TV and Zomato. Mayank holds a B.Tech in Computer Engineering from JIIT Noida.
Two-Way Doors FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




