Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

Thought LeadershipAI Testing

Don't Delete Your Tests - Move Them Up the Stack

Don't Delete Your Tests - Move Them Up the Stack: a startup deleted 811,883 lines of unit tests in one PR. See why the test pyramid breaks with coding agents.

Published on:

This week, a startup deleted 811,883 lines of unit tests in one pull request, and engineers on X couldn't stop arguing about it.

Sherwood Callaway, founder of Sazabi, posted a screenshot of the diff: +166, −811,883. His post passed 202K views and 1.8K likes within hours, and X turned the debate into a trending story titled "Developers Delete 800,000 Lines of Unit Tests Amid AI Coding Shift."

It struck a nerve because every team shipping with coding agents has felt the same tension. Agents write code fast. They also write tests fast, and those tests pile up quicker than anyone can review them. So the question isn't whether Sazabi was reckless. It's whether the way we test still fits the way we now build software.

The Case Against Unit Tests, Taken Seriously

Callaway's argument deserves a fair hearing, because parts of it are right. He offered a few hypotheses for why the deletion made sense. Two of them stand out:

  • Unit tests lock in sloppy code - When an agent writes a weak implementation and a matching unit test, the test protects the weak implementation. Refactoring or deleting it means rewriting the tests too, so bad code stays put.
  • Unit tests slow agents down - Every change triggers more local test runs for the agent and longer CI pipelines. At agent speed, that overhead adds up across hundreds of changes a day.

There's a third problem he didn't need to say out loud. Many agent-written unit tests check that the code does what the code does. They mirror the implementation instead of the intent, so they pass even when the feature is broken for users.

The Catch: The Risk Doesn't Disappear; It Moves

The sharpest reply came from Rails performance consultant Nate Berkopec. He said the idea isn't stupid and is worth experimenting with. But the trade-off is that you need far more testing at higher levels: acceptance, integration and end-to-end testing.

That's the part of the story most hot takes skipped. Deleting 800K lines of unit tests doesn't remove the need to verify software. It moves that job to the layer that checks what users actually experience.

Berkopec also made the point that changes the math: higher-level tests can now be built much faster than before. For years, E2E tests sat at the narrow top of the test pyramid because they were slow to write, brittle to maintain and expensive to run. If AI removes those costs, the pyramid stops making sense.

Why the Testing Pyramid Breaks in the Agent Era

The classic pyramid assumes humans write code slowly and E2E tests are costly. Coding agents flip both assumptions.

When an agent can rewrite a module in minutes, tests coupled to its internals become friction. Tests tied to user behavior stay valid no matter how the code underneath changes. That makes intent-level tests the most durable asset a team owns.

The new shape keeps a thin layer of unit tests for pure logic that rarely changes: pricing rules, parsers, security checks. Most of the weight moves to integration and E2E tests that describe what the product must do, in language a person can read, and an agent can run.

Classic test pyramid vs. agent-era shape

Classic test pyramid vs. agent-era shape

What Agentic E2E Testing Looks Like

The goal is simple: describe how the product should behave, and let an agent turn that into tests that run and keep up as the code changes. That's what we built KaneAI to do.

  • Tests start from intent, not implementation - You write "a returning user can apply a coupon at checkout" in plain English. KaneAI plans the steps and builds the test.
  • Tests survive refactors - Because they target user behavior rather than function signatures, an agent can rewrite the checkout service without touching them, and when the UI itself shifts, KaneAI's self-healing re-anchors the steps for you to review.
  • Tests stay readable - Anyone on the team can review a test written as steps a user takes. That makes slop easy to spot, which is the exact problem Callaway was worried about.
  • Tests run at agent speed - Suites run in parallel on TestMu AI's HyperExecute cloud, so moving weight up the stack doesn't bring back the hour-long CI pipelines teams were trying to escape.

This is how you get Berkopec's "way more testing at higher levels" without hiring a QA team to write it by hand.

Before You Delete Anything: A Playbook

If the Sazabi story tempts you, do it in this order:

  • Map your critical user journeys first - Sign-up, checkout, billing, core workflows. These are what you can't afford to break.
  • Cover those journeys with E2E tests before you remove anything - The safety net goes up before the old one comes down.
  • Sort your unit tests into two piles - keep the ones that guard pure logic and edge cases. Flag the ones that only mirror implementation details.
  • Delete in batches and watch escape rates - Track bugs that reach production for a few weeks after each batch.
  • Give your coding agents the E2E suite as their definition of done - An agent's change ships only when the user journeys still pass.

The Real Lesson

Sazabi didn't prove that testing is dead. It showed that the cheapest place to test has moved. When code is written by agents and rewritten every week, the tests worth keeping describe what users need, not how today's code happens to work.

So don't delete your tests. Move them up the stack, and let an agent carry the weight.

Try KaneAI free and turn your critical user journeys into E2E tests in plain English.

Automate web and mobile tests with KaneAI by TestMu AI

Author

...

Shantanu Wali

Blogs: 9

  • Linkedin

Shantanu Wali is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he owns several product lines across the testing platform, including the Real Device Cloud and the Digital Experience Testing Cloud. He has also contributed significantly to the development and scaling of KaneAI, TestMu AI's flagship GenAI-native testing agent that uses natural language to make software testing faster and more reliable in this AI era. He brings 7+ years of experience across software development and product management, starting as a backend developer at Infosys building solutions for Fortune 500 clients. Shantanu holds an MBA from IIM Calcutta and a B.Tech in Mechanical Engineering.

Reviewer

...

Anmol Gupta

Reviewer

  • Linkedin

Anmol Gupta is Vice President of Product Management at TestMu AI (formerly LambdaTest), driving HyperExecute, the test orchestration cloud that runs and accelerates automated test execution. He led the development of the Unified Test Execution Cloud Platform and now leads a 30-member cross-functional product organization across product lines contributing $7M+ in revenue. He brings over nine years of experience and previously co-founded the SaaS company Timble as CTO, where he grew the team from 5 to 40 and launched an AI KYC platform that processed 600K+ applications in five months while cutting verification time from 12 minutes to under 30 seconds. Anmol holds an MTech and BTech from IIT Delhi.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests