Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIPlaywrightAutomation

Agentic Testing in the End-to-End (E2E) Stack: Fixing Reliability Problems of Playwright MCP Agents

Agentic testing explained: why Playwright MCP and CLI agent runs get slow and costly, and 10 Kane CLI workflows that move agentic E2E tests into CI.

Published on:

Agentic End-to-End Testing (E2E) gets expensive and slow when your coding agent drives the browser itself, because every browser snapshot accumulates in its context. Move the browser loop into a dedicated testing agent and replay what has already passed. Judge every run against written acceptance criteria and keep the evidence. That one change is what moves agentic tests from occasional exploration into your CI gate.

Slack Engineering ran more than 200 agentic End-to-End Testing (E2E) workflows and published the most honest numbers so far. This guide starts with their findings, explains the cause of each, and provides the commands to fix them with Kane CLI, an AI testing agent that runs from your terminal.

TL;DR

Agentic testing is End-to-End Testing (E2E) in which an AI agent receives a goal, works out the steps itself, and is judged against written acceptance criteria. It gets slow and costly when your coding agent drives the browser, so moving that loop to a dedicated testing agent and replaying passed flows fixes both problems.

Why Do Agentic E2E Tests Cost So Much?

  • Re-sent context: The model API is stateless, so every turn re-sends the whole conversation, including every browser snapshot. Slack Engineering measured agentic runs with Playwright MCP and Playwright CLI at about $15-30 each.
  • Delegated browser loop: Kane CLI, the AI testing agent from TestMu AI, drives Chrome in its own process. The coding agent reads one run_end JSON line per test instead of dozens of snapshots.
  • Replay and self-healing: Kane CLI's testmd run command replays cached recordings of flows that already passed and re-authors only the steps that broke, so repeat runs skip fresh reasoning.

Does Agentic Testing Belong Only at the Top of the Testing Pyramid?

  • Three levels: No, agentic testing does not belong only at the top of the testing pyramid. With replay and criteria-based verdicts, it runs as a self-check while coding, a replayed gate on every pull request, and nightly exploration.
  • Pull request gate: Yes, agentic tests can gate pull requests. Kane CLI fails the build through typed exit codes: 1 for a failed test, 2 for a setup error, and 3 for a timeout.

What Is the Difference Between Playwright MCP and Playwright CLI?

  • Playwright MCP: Exposes browser actions as tool calls that return page snapshots to the agent. Slack measured it as the more reliable and faster of the two, and it suits interactive work in one editor session.
  • Playwright CLI: Runs browser commands from the shell one step at a time. In Slack's runs, most of its failures came from login, timing, and session problems rather than reasoning.

What Is Agentic Testing?

Agentic testing is end-to-end testing in which an AI agent receives a goal, such as "reply in a thread and confirm it appears in All Threads", and figures out the steps itself by observing the app, taking actions, and verifying the result. Traditional E2E tests follow a fixed script of clicks and assertions.

Slack summed up the difference well: scripted tests enforce journeys, while agents verify goals. One rule, the design principle Kane CLI is built on, makes agentic tests trustworthy enough for CI: reasoning owns the path, verification owns the verdict. The agent may take any route it likes, but whether the test passed must be decided against written criteria, never by the agent's own account of what it did.

Agentic testing is one branch of AI automation testing, and it also goes by agentic AI testing or agentic QA. It is different from AI agent testing, which checks how an AI agent itself behaves.

What Did Slack Find When It Tested Agents on End-to-End Testing (E2E) Flows?

Slack ran agents through Playwright MCP and the Playwright CLI, and had an agent generate Playwright tests for a simple flow and a complex one, with 20 runs per configuration.

ApproachFailure rate, simple flowFailure rate, complex flowAverage runtime
Agent + Playwright MCP0%~12%~5-8 min
Agent + Playwright CLI~12%~20%~9-11 min
Agent-generated Playwright tests~8%~48%~3 min

They reported agentic runs costing about $15-30 each. Only about one in five runs took the same sequence of actions. Most CLI failures came from login, timing, and session problems rather than reasoning. They concluded that agentic testing belongs at the top of the testing pyramid, for analysis and debugging, not frequent CI runs.

Their numbers are right for how those runs were built. The conclusion changes once you change the build.

Slack's findingRoot cause (from Slack's own analysis)Our answer, with the Kane CLI mechanism
$15-30 per agentic runThe coding agent carries browser snapshots in its own conversation; most tokens are re-sent contextDelegate the browser loop to a dedicated testing agent. The coding agent reads one run_end line (tail -1), so its context barely grows
5-11 minutes per runThe agent re-reasons every step on every runkane-cli testmd run replays cached recordings and only authors steps that need it
Generated Playwright tests failed ~48% on the complex flowBrittle targeting; UI state varianceAdaptive healing by default: smaller replay windows first, then re-authoring the failing steps
Playwright CLI failures were mostly auth, timing, sessionExecution layer, not reasoningPersistent Chrome profiles, secret variables, and exit code 2 for setup errors so infra failures never count as test failures
CLI runs were hard to parallelizeShared session stateOne isolated Chrome per --agent --headless process; testrun run --parallel N; --remote --parallel N on the cloud grid
Only ~20% of runs took the same pathAgents choose their own pathFine, as long as the verdict is independent: acceptance-criteria checkpoints and --final-validation on judge the outcome, not the path
No audit trail described-Every run writes an .evidence pack with steps, screenshots, network and console logs
Scope: single-session web flows-Same commands for native iOS Simulator and Android Emulator apps, and remote devices
Conclusion: agentic testing belongs only at the top of the pyramid, for exploration and debuggingCost and speed made CI use impracticalOnce run replay, agentic tests can gate pull requests. We give a 5-layer stack that uses agents at three levels

Why Is Agentic End-to-End Testing (E2E) So Expensive?

Agentic E2E tests are expensive because the model API is stateless: every turn re-sends the system prompt and the whole conversation, including every browser snapshot taken so far. Slack found most tokens were content the model had already seen. Cost grows with the number of turns and how fast context grows, not with how smart the model is.

The fix is architectural. Don't let your coding agent hold the browser loop.

IN-LOOP (what Slack measured)
Coding agent ──► browser action ──► snapshot back into coding agent's context
     ▲                                                     │
     └──────────── repeat 40-85 turns, context keeps growing ┘

DELEGATED (Kane CLI)
Coding agent ──► kane-cli run "<goal>" --agent --headless
                     │  (testing agent drives Chrome in its own process)
                     ▼
Coding agent ◄── one run_end JSON line: status, summary, final_state, test_url

The coding agent's context grows by one result line per test, not by dozens of snapshots. The browser work happens in a separate process, which also makes it easy to run many tests in parallel. In agent mode, Kane CLI streams NDJSON events and always ends with a single run_end line.

Kane CLI - Testing Agent in Your Terminal

Playwright MCP vs Playwright CLI vs Kane CLI: Which of These AI Testing Tools Should Your Agent Use?

Use Playwright MCP when you want your coding agent to explore a page interactively in the same session. Use Playwright CLI for scripted, step-by-step browser commands. Use Kane CLI when the agent needs a verified pass or fail it can act on, at CI speed, with evidence.

Playwright MCP (agent-driven)Playwright CLI (agent-driven)Kane CLI
How the agent drives itTool calls; page snapshots return into the agent's contextShell commands, one browser step at a timeOne command per goal in plain English
What comes back to the coding agentA snapshot on every stepCommand output on every stepOne run_end JSON line, plus an optional stream of step events
Who decides pass or failThe agent's judgment, unless it also writes assertionsThe agent's judgment, unless it also writes assertionsCriteria you write, checked against the page DOM by default (--assertion-mode dom)
Re-running a flow that passedThe agent reasons through it againThe agent reasons through it againReplays the cached recording; re-authors only steps that broke
When the UI changesThe agent re-reasons from scratchThe agent re-reasons from scratchAdaptive healing, then re-authoring of the failing steps
Setup failure vs test failureSame failure signalSame failure signalSeparate exit codes: 1 failed, 2 setup error, 3 timeout
Parallel runsSupportedHard in Slack's setupOne isolated Chrome per process; testrun run --parallel N locally or --remote on the cloud grid
Audit trailWhatever logs you collectWhatever logs you collect.evidence pack per run
Native mobile appsNot supported (web only)Not supported (web only)iOS Simulator and Android Emulator
Exit to plain codeNot applicableNot applicableExport to Python or JavaScript Playwright

The comparison is about how an agent uses each tool for testing. Playwright itself is excellent, and it's why Kane CLI exports to it. If you are choosing for AI test automation in CI, start with the "Who decides pass or fail" row.

Where Does Agentic Testing Fit in the Testing Pyramid?

Agentic testing belongs at three levels of the testing pyramid, not just the top. With replay and criteria-based verdicts, agents can verify changes while you code, gate pull requests, and explore at night.

  • Unit and integration tests. Deterministic and fast. Agents write them; they don't run them.
  • Agent self-check while coding. Your coding agent verifies the flow it just changed before it says "done".
  • Replayed agentic E2E gate on every pull request. Proven _test.md flows, replay quickly, heal when the UI moves, and fail the build with a typed exit code.
  • Nightly agentic exploration. Open-ended goals with bug detection on, to find what nobody wrote a test for.
  • Production bug reproduction and flaky-test triage. Turn a ticket into a reproducible run with evidence attached.

Keep exported Playwright code for the flows you want fully deterministic forever. Everything else runs as intended.

If your team calls it the test automation pyramid, the split is the same: deterministic tests at the base, agentic checks above them.

Note

Note: Kane CLI is free to install, and local runs are free on a TestMu AI account. Create your free account to try the local workflows below.

10 Kane CLI Workflows for Your Agentic Testing Stack

Every command below comes from the Kane CLI docs. For more recipes, including network, console, and Core Web Vitals checks, see the Kane CLI cookbook.

1. Install Kane CLI and Give Your Coding Agent the Skill

npm install -g @testmuai/kane-cli

# Install the skill for Claude Code, Codex CLI, and Gemini CLI in one go
npx @testmuai/kane-cli-skill

# Agents can't complete browser OAuth, so use basic auth
kane-cli login --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY"
kane-cli whoami

Now ask your agent things like "verify the checkout flow on staging works". The skill builds the command, parses the result, and diagnoses failures from the logs and screenshots.

2. Make Your Coding Agent Verify Its Own Change

kane-cli run "Log in as {{email}}, open Billing, and verify the plan shows Enterprise" \
  --url https://staging.myapp.com \
  --variables-file ./.testmuai/variables/staging.json \
  --agent --headless --timeout 300 --max-steps 50 \
  2>/dev/null | tail -1 | jq '{status, one_liner, reason, test_url}'

The agent gets four fields back, not a page of snapshots. Exit code 0 means it may say "done".

For a quick check of a page that's already open, attach to that Chrome over CDP and judge conditions without taking any actions:

kane-cli run --analyzer-only --agent \
  --cdp-endpoint "$CDP_URL" \
  --condition "The pricing table shows three plans" \
  --condition "The Pro plan shows a monthly price"

Each verdict comes back in condition_results on run_end.

3. Gate Pull Requests With Replayed Agentic Tests

Keep proven flows as _test.md files in your repo. testmd run replays cached recordings and authors fresh steps only when needed. When a step breaks, adaptive healing retries smaller replay windows and then re-authors only the failing steps, which is what makes these self-healing tests.

# One test
kane-cli testmd run tests/checkout_test.md --headless --agent

# The smoke suite, 4 workers, stop at the first failure
kane-cli testrun run --match 'tests/e2e/.*' --tags smoke \
  --parallel 4 --on-failure fail-fast --headless \
  --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY"

In GitHub Actions:

name: Agentic E2E gate
on: [pull_request]

jobs:
  kane:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: '20'
      - uses: browser-actions/setup-chrome@v1
      - run: npm install -g @testmuai/kane-cli
      - name: Run smoke suite
        env:
          LT_USERNAME: ${{ secrets.LT_USERNAME }}
          LT_ACCESS_KEY: ${{ secrets.LT_ACCESS_KEY }}
        run: |
          kane-cli testrun run --tags smoke --parallel 4 \
            --on-failure fail-fast --headless \
            --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY"
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: kane-evidence
          path: |
            .testmuai/evidence/*.evidence
            ~/.testmuai/kaneai/sessions/

4. Run the Suite on the Cloud Grid When the Runner Can't

kane-cli plugin install remote-execution

kane-cli testrun run --match 'tests/web/.*' --remote --parallel 4 \
  --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY"

The grid supplies the browsers, so the CI runner needs no Chrome. Remote runs need a plan with HyperExecute macOS runners. The recordings and the evidence pack return to the checkout.

5. Explore at Night and Let the Agent Hunt Bugs

kane-cli run "Explore account settings: change the display name, upload an avatar, \
  switch the language to French and back, and report anything broken" \
  --url https://staging.myapp.com \
  --bug-detection continue \
  --agent --headless --max-steps 80 --timeout 900

continue records each confirmed bug and keeps going, so one run can surface several issues.

6. Reproduce a Production Bug From a Ticket

kane-cli run "Add two items to the cart, apply code SAVE10, remove one item, \
  and verify the discount is recalculated on the remaining item" \
  --url https://staging.myapp.com \
  --final-validation on --agent --headless --timeout 300

Paste the test_url and the .evidence pack into the ticket. Anyone can replay exactly what happened. The pack follows an open format for test run evidence, with its own CLI to validate it.

7. Separate Flaky Tests From Flaky Infrastructure

Run the same goal in parallel and read the exit codes. 1 means the product failed; 2 means setup failed (auth, Chrome, config); 3 means timeout.

for i in 1 2 3 4 5; do
  kane-cli run "Search for 'quarterly report' and open the first result" \
    --url https://staging.myapp.com --agent --headless --timeout 180 \
    > run_$i.ndjson 2>/dev/null &
done
wait
for f in run_*.ndjson; do tail -1 "$f" | jq -r '.status'; done | sort | uniq -c

Five passes points to the old script, not the product, as the source of the flakiness. Mixed results with code 1 mean a real, intermittent product bug.

8. Export to Playwright for Flows You Want Frozen

kane-cli run "Sign up with a new email and verify the welcome screen" \
  --url https://staging.myapp.com --code-export --code-language javascript \
  --agent --headless

# Or regenerate code from an existing test.md recording
kane-cli testmd export tests/signup_test.md --language javascript

Use this Playwright AI handoff when a flow should stop changing: Kane CLI authors it once, and the exported Playwright test lives in your repo like any other.

9. Test Native Mobile Apps With the Same Commands

kane-cli doctor --target emulator --install
kane-cli devices list --target emulator

kane-cli run "Log in, open Orders, and verify the latest order shows Delivered" \
  --target emulator --device-name "Pixel 7" --os-version 14 \
  --app ./app-debug.apk --agent

Runs on macOS with Apple Silicon. Use --remote with testrun run to run on cloud devices instead.

10. Extract Data to Feed the Next Step

kane-cli run "Open the pricing page and store the Pro monthly price as 'pro_price'" \
  --url https://myapp.com --agent --headless 2>/dev/null \
  | tail -1 | jq -r '.final_state.pro_price'

Your agent can compare the value with the pricing config it just edited.

Get Kane CLI certified for free with TestMu AI

How Do You Keep Agentic Test Costs Under Control?

Most of the spend in AI E2E testing comes from re-sent context and repeated reasoning.

  • Delegate the loop. The coding agent should read results, not snapshots.
  • Replay what has passed. Author once with run, keep it as _test.md, replay with testmd run.
  • Cap every run. Always set --max-steps and --timeout.
  • Measure. run_end includes token_usage; track it per test like a performance budget.
  • Fan out, don't chain. Run independent goals as parallel processes.
  • Save exploration for nightly. Open-ended goals cost the most; run them on a schedule, not every commit.

How Do You Write a Good Agentic Test Objective?

A good objective states the goal and the outcome to verify. A bad one lists clicks or leaves the outcome vague.

Weak objectiveStrong objective
"Test checkout""Add a medium blue t-shirt to the cart, check out as guest with {{card}}, and verify the confirmation page shows an order number"
"Click Settings then Profile then Save""Change the display name to {{name}} and verify it appears in the header"
"Check search works""Search for 'invoice 2026' and verify at least one result opens a PDF"

Use {{variables}} for test data and secrets. Never paste credentials into the objective.

How Do You Verify These Claims on Your Own App?

Run Slack's experiment yourself. Pick one real flow and run it 20 times as a fresh agentic run, then 20 times as a replay. Compare pass rate, duration, and tokens. The numbers settle the argument for your app, not ours or Slack's.

The same comparison is a fair way to judge any AI QA testing claim, including the ones in this guide.

# 1. Twenty fresh agentic runs
OBJ="Create a channel, post a message, reply in a thread, and verify the reply appears in All Threads"
for i in $(seq 1 20); do
  kane-cli run "$OBJ" --url https://staging.myapp.com \
    --agent --headless --timeout 600 --max-steps 60 2>/dev/null \
    | tail -1 | jq -c '{status, duration, token_usage}' >> fresh.ndjson
done
jq -s '{runs: length, passed: (map(select(.status=="passed")) | length), avg_seconds: (map(.duration) | add / length)}' fresh.ndjson

# 2. Save the flow as tests/thread_reply_test.md, then twenty replays
for i in $(seq 1 20); do
  start=$(date +%s)
  kane-cli testmd run tests/thread_reply_test.md --agent --headless > /dev/null 2>&1
  echo "{\"exit\": $?, \"seconds\": $(( $(date +%s) - start ))}" >> replay.ndjson
done
jq -s '{runs: length, passed: (map(select(.exit==0)) | length), avg_seconds: (map(.seconds) | add / length)}' replay.ndjson

See Running test.md files for how to save a flow as a _test.md file.

Objections, Answered

"Isn't delegating the loop just moving the tokens somewhere else?" The first authoring run still uses a model; that cost doesn't disappear. Two things change. The coding agent's context stops growing with browser snapshots, which Slack identified as the main cost driver. And replays of steps that have already passed don't need fresh reasoning. Every run_end reports token_usage, so you can check this on your own flows.

"The verifier is AI too. Why trust its verdict?" You write the acceptance criteria, not the model. By default, assertions read the page DOM (--assertion-mode dom) and use vision only as a fallback. Every verdict ships with an .evidence pack anyone can open and replay, so a wrong verdict is visible and checkable.

"What about apps that really are nondeterministic?" Two of the usual causes of flaky tests are gone by construction: there are no hand-written locators to break and no hand-tuned waits to guess wrong. For real product nondeterminism, bug detection classifies what it sees and reports a confidence signal. Kane's own docs call this partial, and it is: no tool fully solves nondeterministic app state.

"Am I locked in?" No. _test.md files live in your repo as markdown, and any flow exports to Python or JavaScript Playwright with --code-export or kane-cli testmd export. You can leave with working code.

"Does my coding agent need to change?" No. The skill is a plain Markdown file for Claude Code, Codex CLI and Gemini CLI, and any agent that can run a shell command can call kane-cli run --agent and read the last line. For agent-specific setup, see the guides on testing apps built with Claude Code and testing apps built with Cursor.

"Is it safe with real credentials?" Pass secrets as variables marked "secret": true or from your CI secret store, never inside the objective. Local runs use your own Chrome on your machine.

What Doesn't Kane CLI Do?

  • Local mobile runs need macOS on Apple Silicon. Other machines can run mobile suites on the cloud grid with testrun run --remote.
  • One objective per run invocation. Parallelism comes from separate processes or testrun run --parallel N.
  • No mid-run questions in CI. ask_user is turned off when there's no terminal, so flows that need a human for OTPs or CAPTCHAs should use test-mode bypasses in CI.
  • Cloud runs need a plan with HyperExecute macOS runners.
  • It needs a TestMu AI account. The free tier is enough to try the local workflows above.

When Should You Still Use Playwright MCP Directly?

Use Playwright MCP directly for interactive work inside one editor session: poking at a page, inspecting the DOM or building a selector by hand. For anything you want to repeat, gate a build on or hand to a teammate, use a testing agent that returns a verdict and evidence.

For MCP testing that should end in a verified pass or fail rather than raw MCP primitives, see how Kane CLI works as a layer above browser automation MCP.

Note

Note: Install Kane CLI and claim 10,000 free credits.

Author

...

Shantanu Wali

Blogs: 9

  • Linkedin

Shantanu Wali is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he owns several product lines across the testing platform, including the Real Device Cloud and the Digital Experience Testing Cloud. He has also contributed significantly to the development and scaling of KaneAI, TestMu AI's flagship GenAI-native testing agent that uses natural language to make software testing faster and more reliable in this AI era. He brings 7+ years of experience across software development and product management, starting as a backend developer at Infosys building solutions for Fortune 500 clients. Shantanu holds an MBA from IIM Calcutta and a B.Tech in Mechanical Engineering.

Reviewer

...

Anmol Gupta

Reviewer

  • Linkedin

Anmol Gupta is Vice President of Product Management at TestMu AI (formerly LambdaTest), driving HyperExecute, the test orchestration cloud that runs and accelerates automated test execution. He led the development of the Unified Test Execution Cloud Platform and now leads a 30-member cross-functional product organization across product lines contributing $7M+ in revenue. He brings over nine years of experience and previously co-founded the SaaS company Timble as CTO, where he grew the team from 5 to 40 and launched an AI KYC platform that processed 600K+ applications in five months while cutting verification time from 12 minutes to under 30 seconds. Anmol holds an MTech and BTech from IIT Delhi.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Agentic Testing FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests