Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- End-to-End (E2E) Testing With an AI Agent: Should Cursor Drive Playwright MCP or Kane CLI?
End-to-End (E2E) Testing With an AI Agent: Should Cursor Drive Playwright MCP or Kane CLI?
Write and Run your end to end (e2e) tests now with Kane CLI. See how it compares with Playwright MCP inside Cursor on context cost, approvals, pass or fail verdicts, CI exit codes, and replays that cost no credits.
Published on:
Cursor has just finished the change to your cart page. The types check and the unit tests pass, and you want one more thing before you open the pull request: someone to load the page in a real browser, add a product, and confirm the cart shows it.
Two tools can give Cursor that browser, and they split the work very differently. With Playwright MCP, Cursor's own model clicks through the page. With Kane CLI, Cursor hands the whole check to a separate browser agent and reads back a verdict. I ran the same end-to-end flow both ways and measured what each one would put back into Cursor's context.
TL;DR
Kane CLI vs Playwright MCP for end-to-end testing comes down to who drives the browser. With Playwright MCP, Cursor's own model calls browser tools one step at a time and decides whether the test passed. With Kane CLI, Cursor hands one natural-language objective to a separate browser agent and gets back a pass or fail verdict with evidence.
- Best for exploring a page while coding: Playwright MCP - Cursor reads live accessibility snapshots, console messages, and network requests through Playwright MCP's 25 default tools, with no account or credits needed.
- Best for verifying a finished feature: Kane CLI - one terminal command returns a run_end result, an exit code of 0, 1, 2, or 3, and an evidence pack with step screenshots and a network log.
- Context cost on a five-step flow: Did Playwright MCP put more into Cursor's context than Kane CLI? - Yes. Playwright MCP returned 42,333 tokens with full snapshots or 4,076 with browser_find, plus 4,396 tokens of tool definitions, against 2,056 from Kane CLI.
- Approval checks in Cursor: Does the Playwright MCP flow need more approval checks? - Yes. The Playwright MCP flow took 12 MCP tool calls, and Cursor's Auto-review and Allowlist modes check each one, by classifier or by prompt, while Kane CLI ran as one terminal command.
- Rerunning the same check: Does a Kane CLI rerun cost credits again? - No. A Kane CLI run saved with --name replayed 9 cached actions in 23 seconds and consumed no credits, while a Playwright MCP rerun means Cursor drives every step again.
What Does Each One Actually Do?
Playwright MCP is a browser tool server for an AI agent. Its Playwright MCP README describes a server that lets LLMs "interact with web pages through structured accessibility snapshots, bypassing the need for screenshots or visually-tuned models." Its tools, such as browser_click, take an "exact target element reference from the page snapshot", so Cursor's model reads a snapshot first and then acts on what it found.
Kane CLI is a browser agent rather than a tool server. Cursor writes a natural-language objective into a kane-cli run command; Kane CLI drives a real Chrome instance through the Chrome DevTools Protocol with its own model, checks the outcome against evidence, and returns one structured result.
That difference decides everything downstream, including how much Cursor reads and who calls the result. If the server itself is new to you, the guide to the Playwright MCP server covers its tools before any of this comparison matters.
How Do the Setups Compare?
Cursor's MCP docs show how to add an MCP server for one project in .cursor/mcp.json, or for every project in ~/.cursor/mcp.json.
Cursor's Run Modes docs say they "control how the Cursor agent runs tool calls, and when Cursor interrupts you for approval." Auto-review "applies to shell, MCP, and Fetch tool calls" and sends calls that are not allowlisted to a classifier, while in Allowlist mode only "Actions in your allowlist run without approval." Cursor 3.6 shipped Auto-review "as the recommended default," and only Run Everything lets "Every tool call" run automatically.
The Playwright MCP entry, as its README gives it, then the Kane CLI setup and the command Cursor ends up running:
# Playwright MCP: .cursor/mcp.json
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
# Kane CLI: install, sign in with an access key, check the account
npm install -g @testmuai/kane-cli
kane-cli login --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY"
kane-cli whoami
# then, in Cursor's agent chat:
# "Use Kane CLI to verify the cart flow: testmuai.com/kane-cli/agents.md"
# the command Cursor runs, in agent mode for NDJSON output
kane-cli run "Search for 'iPhone' using the search box, open the first iPhone product in the results, click Add to Cart, then open the shopping cart and assert the cart contains a product named 'iPhone'" \
--url https://ecommerce-playground.lambdatest.io/ \
--agent --headless --name iphone-cart --timeout 600| Dimension | Playwright MCP | Kane CLI |
|---|---|---|
| Who drives the browser | Cursor's model, one tool call per action | Kane CLI's own agent; Cursor issues one command |
| How the page is read | Accessibility snapshots, with no vision model | Vision plus DOM, network, and console checks |
| Who decides pass or fail | Cursor's model, unless you enable the opt-in testing tools | Kane CLI, through the status and exit code in its result |
| What Cursor gets back | Tool responses, page snapshots, and generated Playwright code per action | NDJSON events ending in one run_end object |
| Approval checks in Auto-review or Allowlist mode | One per MCP tool call that is not allowlisted | At most one, for the terminal command |
| Account and cost | Apache-2.0; no account; cost is Cursor's model tokens | Apache-2.0 package; TestMu AI account; credits when a run is authored |
The Kane CLI with AI coding agents docs cover the Cursor setup step by step, including why an agent should sign in with an access key rather than a browser login.
Who Owns the Selectors and the Maintenance?
Playwright MCP's element references live only as long as the snapshot that produced them. In my run, the search box was e216 on the home page, and after the next navigation the references came back with a new prefix, such as f1e495 for the product link. Nothing about how Cursor found an element survives the session.
What does survive is code. Each click and type response included a "Ran Playwright code" block with a role-based locator, and the opt-in browser_generate_locator tool, enabled with --caps=testing, turns an element into a locator, so Cursor can write a Playwright spec from the session. From then on, the locators are yours to keep current through every redesign.
Kane CLI keeps the objective instead. The run saved with --name became a 303-byte iphone-cart_test.md holding the sentence, plus a cached recording that replays without a model.
- Playwright MCP's maintenance surface - The Playwright code Cursor writes from the session: locators, waits, and assertions, all versioned with the app and all needing an update when the UI changes shape.
- Kane CLI's maintenance surface - The wording of the objective. A cached step that no longer matches the page needs re-authoring, and a vague objective can be satisfied by the wrong screen.
- What a UI change looks like in Kane CLI - In the Claude Code walkthrough in the Kane CLI docs, a replay failed after a button was renamed from Create to Add project. The agent edited that one step so it re-authored against the new UI.
Neither surface is free. The question is which one your team can keep accurate as the product moves.
Which One Fits an AI Coding Agent's Loop?
This is where the two diverge most, so I measured it. The flow ran on TestMu AI's Ecommerce Playground: open the store, search for "iPhone", open the first result, add it to the cart, and confirm the cart lists an iPhone.
- Playwright MCP 0.0.83 - I connected through the official MCP SDK client, the same protocol Cursor uses, and made the tool calls a Cursor agent would make, without a model in between. That isolates what the tools themselves return.
- Two reading strategies - In the first, the agent reads the page with a full
browser_snapshotbefore each action. In the second, it usesbrowser_find, which searches the snapshot and returns only matching nodes. - Why the second read is needed - In 0.0.83, action tools such as
browser_navigateandbrowser_clicksave the snapshot to a file and return only a link, so the agent has to ask for the page again before its next action. - Kane CLI in agent mode - I ran the same flow as one objective with
--agent --headless, then replayed the saved test. - Token counts - Every response was counted with the o200k_base tokenizer. Cursor's models tokenize slightly differently, and the counts leave out Cursor's own reasoning, which grows with each of the 12 Playwright MCP turns.
| Measure | Playwright MCP, full snapshots | Playwright MCP, browser_find | Kane CLI, first run | Kane CLI, replay |
|---|---|---|---|---|
| Setup text loaded | 4,396 tokens of tool definitions | 4,396 tokens of tool definitions | 1,784 tokens, if Cursor reads the hosted skill | Same skill, already read |
| Calls from Cursor | 12 tool calls | 12 tool calls | 1 command | 1 command |
| Text returned to Cursor | 42,333 tokens | 4,076 tokens | 2,056 tokens in 24 NDJSON lines | 3,575 tokens in 53 NDJSON lines |
| Final result line only | None; Cursor decides | None; Cursor decides | 764 tokens | 65 tokens |
| Browser time | 15 seconds, plus Cursor's thinking on 12 turns | 15 seconds, plus Cursor's thinking on 12 turns | 96.7 seconds, reasoning included | 23 seconds |
| Kane CLI credits | None | None | 22.8 credits | None |
All four runs confirmed an iPhone in the cart. The per-call log shows where the tokens went:
- Snapshots are the cost - With full snapshots, four page reads made up 41,385 of the 42,333 tokens. The store's home page alone was 20,116 tokens, a 1,360-line accessibility tree with 1,156 element references.
- browser_find helps, if the agent knows what to look for - Searching for "Search For Products" or "Add to Cart" cut the reads to a few hundred tokens each. The search-results page still returned 2,241 tokens, because "iPhone" matched every product card.
- Kane CLI's result is small and structured - Cursor can read only the final run_end line, which carries the status, a summary, and the values Kane CLI extracted. Kane CLI's own model work is billed as credits, outside Cursor's context.
- Approval checks add up - The Playwright MCP flow made 12 MCP tool calls, which means 12 prompts under an empty allowlist or 12 classifier reviews under Auto-review. The Kane CLI command needed at most one, and only Run Everything skips the checks.
Microsoft makes the same point in its own README, recommending that coding agents favor CLI workflows "because CLI invocations are more token-efficient: they avoid loading large tool schemas and verbose accessibility trees into the model context." The broader trade-off is covered in MCP vs CLI.
How Do They Behave in CI?
A pipeline needs a verdict it can gate on, and the two tools produce that very differently.
With Playwright MCP, the verdict is Cursor's judgment. The server does ship assertion tools, such as browser_verify_text_visible and browser_verify_element_visible, but they are opt-in through --caps=testing. Without them, "the cart has an iPhone" means Cursor read a snapshot and concluded so, and nothing produces an exit code a pipeline can gate on. For CI, the usual path is to have Cursor write a Playwright spec and run that with the Playwright test runner.
Kane CLI makes the verdict part of the output. Here is the run_end line from the first run, trimmed:
{
"type": "run_end",
"status": "passed",
"one_liner": "added an iPhone to the cart on LambdaTest Ecommerce Playground",
"final_state": {
"url": "https://ecommerce-playground.lambdatest.io/index.php?route=checkout/cart",
"cart_contains_iphone": "true"
},
"reason": "Objective completed",
"duration": 96.7,
"credits_consumed": 22.785712500000002,
"result_code": 100,
"reason_code": "success.complete"
}- The full line also shows how the assertion was checked: a DOM check that the cart row's text equals "iPhone", with the locator it used.
- The process exited with code 0. The Kane CLI reference lists 1 for a failed assertion, 2 for an error such as a Chrome or auth failure, and 3 for a timeout, so CI can tell a broken feature from a broken environment.
- Combine --agent with --headless on any runner without a display server, and sign in with an access key, since an OAuth login needs a browser window the runner does not have.
- The run sealed an evidence pack on disk with 9 step screenshots, a HAR network log, console logs, and the test definition.
Note: Kane CLI runs a natural-language objective in a real Chrome browser and returns pass or fail with an evidence pack your CI can gate on. Start free with TestMu AI and run your first objective from Cursor's terminal.
Should You Run Both Together?
For most teams using Cursor, yes. Playwright MCP's README says MCP "remains relevant for specialized agentic loops that benefit from persistent state, rich introspection, and iterative reasoning over page structure", and that describes real jobs inside the editor:
- Debugging while you code - Cursor can open the page, read console messages and network requests, and change the code in the same session.
- Drafting Playwright tests - The per-action code blocks, plus
browser_generate_locatorwhen--caps=testingis on, give Cursor real locators for a spec your team already runs. - No account, no credits - It runs on Node.js 18 or newer, so a quick look at a page costs only Cursor's tokens.
- Fine-grained control - Opt-in capability groups add storage, network mocking, tracing, video, and PDF tools when a task needs them.
The README is also direct that "Playwright MCP is not a security boundary", so treat what it reads as untrusted input to Cursor. A split that works puts each tool where its cost profile fits: Playwright MCP while Cursor builds, Kane CLI when the feature is done and needs a verdict.
Kane CLI's saved test is what makes the second half cheap. The replay of the cart flow reported this:
kane-cli testmd run .testmuai/tests/iphone-cart_test.md --agent --headless
# from the NDJSON stream
"replay_started" "replaying 9 actions"
"replay completed" duration 13.69
"replay_decisions": 1, "author_decisions": 0
{"type":"test_md_done","overall_status":"passed","duration_s":23}Zero author decisions means no model was asked to work anything out, which matches the Kane CLI Test.md docs: after the first run, steps replay from cache "with no LLM cost." The first run also produced a Playwright Python export in which four of the five locators were role-based, so a flow that starts as an objective can still end up in a Playwright suite. The file format behind that round trip is covered in Test.md, and the side-by-side capability table lives on the Playwright MCP alternative page.
Which Should You Choose?
Match the tool to the job in front of you rather than to a feature count:
| Your situation | Start with | Why |
|---|---|---|
| Inspecting a page while Cursor writes the code | Playwright MCP | Live snapshots, console, and network in the same session, at no cost beyond tokens. |
| Confirming a finished feature before the PR | Kane CLI | One command, one verdict, and evidence you can attach, with little added to Cursor's context. |
| Gating a pipeline on an e2e check | Kane CLI in agent mode | Distinct exit codes separate a failed assertion from an environment error. |
| Growing a Playwright suite your team maintains | Playwright MCP to draft, the Playwright runner to execute | The session yields real locators and code that land in a framework you already run. |
| Rerunning the same check on every build | Kane CLI with a saved test | Cached replays ran without spending credits or model turns. |
If you are weighing more than these two, Playwright MCP alternatives compares five options, and Playwright CLI vs Kane CLI covers the same choice for Playwright's terminal tool.
Conclusion
Take the one journey your team would least like to ship broken and run it both ways from Cursor this week: once with Playwright MCP connected, once as a Kane CLI objective. Compare what each put into Cursor's context and which answer you would trust in a pull request.
Start with npm install -g @testmuai/kane-cli and the Kane CLI quickstart, which gets a first objective running in a real Chrome browser.
Author
Srinivasan Sekar is Director of Engineering at TestMu AI (formerly LambdaTest), where he leads engineering and open-source initiatives behind the Selenium and Appium automation grid and owns TestMu AI's MCP Server. A committer to Appium and a contributor to Selenium, WebdriverIO, Taiko, and AppiumTestDistribution, he brings over 15 years of experience in quality engineering and open-source technologies. He is the author of the Apress book 'The MCP Standard: A Developer's Guide to Building Universal AI Tools with the Model Context Protocol,' a Certified Kubernetes and Cloud Native Associate, and an international conference speaker. Before TestMu AI he spent over eight years at Thoughtworks as a Principal Consultant and Quality Architect. Srinivasan holds a B.Tech in Information Technology from Anna University.
Reviewer
Mayank Bhola is Co-Founder and Head of Products at TestMu AI (formerly LambdaTest), where he leads the entire product portfolio across KaneAI, Kane CLI, HyperExecute, SmartUI, the Real Device Cloud, Accessibility, and other software testing product lines. As an early Lead Architect he designed and built the company's flagship Tunnel technology from scratch, created the React-based automation platform, and architected the data-intensive pipelines and FAAS services that scale it. He brings more than 10 years of experience in software development and product engineering, with earlier roles as Head of Technology at Juggernaut Books and Senior Software Engineer at PressPlay TV and Zomato. Mayank holds a B.Tech in Computer Engineering from JIIT Noida.
Kane CLI vs Playwright MCP FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




