Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AutomationAI Testing

e2e by TesterArmy: Open-Source Agentic E2E Testing Explained

e2e by TesterArmy is an open-source agentic E2E testing framework. See how its tests, record and replay, models, and CI work, and how to run it on TestMu AI.

Published on:

Scripted end-to-end tests break when a selector or a flow changes, and fully agentic tests ask a model to rediscover every step on every run. e2e by TesterArmy sits between the two: an agent handles the steps you describe in plain English, and ordinary locators and assertions check the result.

TesterArmy has open-sourced it under Apache-2.0, and its tests run locally, in CI, and on hosted browsers, including TestMu AI cloud browsers through the web engine's CDP connection.

TL;DR

e2e by TesterArmy is an open-source, Apache-2.0 end-to-end testing framework for web and mobile apps in which a test step can be a plain-English goal that an agent completes, checked by locators and assertions in the same test. Verified agent steps replay from cache without model calls, and any AI SDK model can drive the agent.

  • Agent steps: Does every e2e test need an AI model? No. Tests built only from locators and assertions run without one; agent.act and agent.assert steps need an AI SDK model.
  • Record and replay: Do replayed agent steps call the model? No. After a later check verifies an agent.act step, e2e replays its recorded actions on the next run without a model call.
  • Mobile apps: Does e2e support mobile apps? Yes. The @e2e-dev/mobile engine runs the same tests on iOS simulators and Android emulators through agent-device.
  • Migration: Can you migrate Playwright or Cypress tests to e2e? Yes. The e2e docs include migration guides from Playwright, Cypress, Selenium, Detox, and Maestro.
  • Cloud runs: Can e2e tests run on TestMu AI? Yes, for web tests. The web engine attaches to TestMu AI Automation Cloud Chrome sessions over the Chrome DevTools Protocol, one session per worker.

What Is e2e by TesterArmy?

e2e is an open-source end-to-end testing framework for web and mobile apps, built by TesterArmy and licensed under Apache-2.0. Its GitHub README describes the model: you describe a goal in natural language, an agent interacts with the app to complete it, and locators and assertions in the same test check exact results.

Running npx e2e init scaffolds a project. The wizard picks an engine and a model provider, including None for tests without AI, writes e2e.config.ts and an example test, and installs a skill that coding agents can use to write and run tests.

Youtube thumbnail

How Does an e2e Test Work?

An e2e test combines fixtures: app opens the app, agent takes plain-English steps, and screen finds elements for exact assertions.

The agent steps docs define a step as one agent.act, agent.assert, agent.waitFor, or agent.extract call, and every step has a deadline and a model-call budget.

This test gives the agent a goal, then checks the outcome with a locator:

// tests/agent-form.e2e.ts
import { test, expect } from 'e2e';

test('an agent fills the form from a plain-English goal', async ({ app, agent, screen }) => {
  await app.open('/selenium-playground/simple-form-demo/');
  await agent.act('enter "Agent run on TestMu AI" in the single message field and submit it with Get Checked Value');
  await expect(screen.getByText('Agent run on TestMu AI', { exact: true })).toBeVisible();
});
  • agent.act - completes a goal such as filling and submitting a form. The model reads a text snapshot of the screen and asks for actions, which the runner performs and records.
  • agent.assert - a model judgment about what is on screen, which sees the instruction and the current screen but not the acting agent's summaries.
  • screen and expect - deterministic locators such as getByRole, getByText, and getByPlaceholder, with matchers such as toBeVisible, that check exact results without a model.

The snapshot the model reads leaves out raw HTML, cookies, headers, and environment values, masks password fields, and replaces configured secrets with a placeholder. A test with no agent steps never calls a model, so existing locator checks can sit in the same suite as agent steps.

How Does Record and Replay Cut Model Calls?

The replay cache records the actions in an agent.act() step after a later check verifies the result. On the next run, the runner replays those actions without calling the model, and if a control or the expected end state no longer matches, the agent takes over from the current screen.

The first run of the test above spends model calls while the agent works out the steps:

   ✓ an agent fills the form from a plain-English goal > agent.act "enter "Agent run on TestMu AI" ..." 16.20s · 3 model calls
 ✓  testmu-chrome  tests/agent-form.e2e.ts (1 test) 29.14s ai 28.4k tokens

 Test Files  1 passed (1)
      Tests  1 passed (1)
         AI  28.4k tokens · 3 model calls · google.generative-ai/gemini-2.5-flash
      Cache  1 missed

The second run replays the recorded actions with no model call:

   ✓ an agent fills the form from a plain-English goal > agent.act "enter "Agent run on TestMu AI" ..." 6.11s
 ✓  testmu-chrome  tests/agent-form.e2e.ts (1 test) 19.73s

 Test Files  1 passed (1)
      Tests  1 passed (1)
      Cache  1 replayed
  • Matching by meaning - replay finds each recorded control by its role, name, test id, and surrounding context, and ignores text that changes with the data, such as counts, dates, and ids.
  • Live judgments - agent.assert, agent.waitFor, and agent.extract always run live, so checks never come from a recording.
  • Run summary - each cached step is reported as replayed, handed off, or missed, and npx e2e run --no-cache rules the cache out while debugging.

Replay is what makes agent tests affordable on every pull request: model calls happen when the app changes. For a rollout plan built around steps like these, see the agentic E2E testing playbook.

Which Models, Engines, and Platforms Does e2e Support?

Agent steps use the model set in e2e.config.ts, and there is no default. The models docs cover a gateway such as Vercel AI Gateway or OpenRouter, a provider API key, or a local or self-hosted server, and advise a model that supports tool calls and images.

Subscriptions work without a separate key. The subscriptions guide signs in with npx e2e login openai for ChatGPT Plus or Pro, github-copilot for GitHub Copilot, and spacexai for SuperGrok or X Premium+. Claude subscriptions are not supported, though Claude works through an API provider.

Engines and add-ons ship as separate npm packages:

PackageWhat it does
e2eSDK, runner, and CLI
@e2e-dev/webBrowser engine built on Playwright: Chromium, Firefox, and WebKit
@e2e-dev/mobileiOS simulators and Android emulators, built on agent-device
@e2e-dev/kernel and @e2e-dev/easHosted Kernel browsers for the web engine and hosted EAS simulators and emulators for the mobile engine
@e2e-dev/githubPull request comment reporter for GitHub Actions

Device tests use the same agent, screen, and expect APIs as browser tests. The mobile docs require Xcode with an iOS simulator runtime or the Android SDK with an emulator, checked with npx agent-device doctor.

The API testing docs show API tests written with fetch and schema matchers, in the same files and the same run as UI tests.

For existing suites, the e2e docs include migration guides from Playwright, Cypress, Selenium, Detox, and Maestro.

How Do You Run e2e in CI?

Deterministic and agent tests run in the same CI job. e2e's CI guide installs browsers with npx playwright install chromium --with-deps, runs npx e2e run --reporter list,junit, and uploads .e2e/report.json and junit.xml as artifacts.

  • Pull request comments - @e2e-dev/github turns each run into one comment on the pull request, edited in place on reruns, with the same text in the job summary.
  • Failure pages - a failed test gets a page under .e2e/failures/ with what was expected and observed, the steps, and the accessibility tree, plus a screenshot and an optional Playwright trace under .e2e/artifacts/.
  • Faster loops - npx e2e run --last-failed reruns what failed, and --shard 2/3 splits the suite across CI jobs.

Who Verifies the Agents Running Your Tests?

Once an AI agent is part of test execution, its own account of what it did cannot be the evidence. e2e builds that rule in: an agent.assert judgment does not see the acting agent's summaries, and only an agent.act step that a later check verifies is recorded for replay.

Agents your team builds need the same treatment once they act on real systems. TestMu AI Agent Assurance tests how your agents actually behave across workflows, tools, and actions: it invokes the agent for real and grades each criterion against observed evidence, such as files changed, artifacts produced, and tool calls, instead of the agent's summary of its work. It is pre-alpha and runs from the terminal as rook.

For coding agents, Kane CLI runs a plain-English browser check against the running app from the terminal before you open a pull request, and returns a pass or fail backed by an evidence pack. In agent mode it emits NDJSON, so a coding agent or a CI job can read the verdict. Other agent-driven options are compared in the roundup of agentic AI testing tools.

Kane CLI - Testing Agent in Your Terminal

How Do You Run e2e on TestMu AI?

e2e's web engine can attach to a remote browser instead of launching one. Its web engine reference documents a connect option that needs Chromium and calls its cdpEndpoint resolver once per worker.

TestMu AI Automation Cloud provides that browser through the CDP endpoint in its Puppeteer testing docs, the same remote-browser pattern as a Playwright grid:

  • Prerequisites - a TestMu AI account, with the username and access key from the Access Key button on the Automation Dashboard exported as LT_USERNAME and LT_ACCESS_KEY.
  • Connect - set connect.cdpEndpoint on the web target to TestMu AI's CDP endpoint, passing your capabilities, as in the config below.
  • Run - run npx e2e run as usual. Each worker opens its own cloud Chrome session, so --workers sets how many tests run in parallel.
  • Review - open the build in Automation Cloud for each session's video and command logs, with network and console logs when the capabilities turn them on. Pass or fail stays in e2e's report.
// e2e.config.ts
import type { E2EConfig } from 'e2e';
import { web } from '@e2e-dev/web';

const capabilities = {
  browserName: 'Chrome',
  browserVersion: 'latest',
  'LT:Options': {
    platform: 'Windows 11',
    build: 'e2e on TestMu AI',
    name: 'e2e suite',
    video: true,
    network: true,
    console: true,
    user: process.env.LT_USERNAME,
    accessKey: process.env.LT_ACCESS_KEY,
  },
};

export default {
  targets: [
    {
      name: 'testmu-chrome',
      engine: web({
        connect: {
          cdpEndpoint: async () =>
            `wss://cdp.lambdatest.com/puppeteer?capabilities=${encodeURIComponent(JSON.stringify(capabilities))}`,
        },
      }),
      app: { url: 'https://www.testmuai.com' },
    },
  ],
} satisfies E2EConfig;

The connection is Chromium-only, so Firefox and WebKit targets keep launching locally, and the web engine's README notes that connect mode does not support headers, basicAuth, userAgent, context reset, or session state capture and restore.

Note

Note: Run your e2e web suite on TestMu AI cloud Chrome, with parallel sessions and a video of every run. Try TestMu AI free!

How Do You Get Started With e2e?

Run npx e2e init, pick the web engine, and write one test that pairs an agent.act step with a locator assertion. Run it twice to see the replay, then add the TestMu AI target above when you need parallel cloud browsers, storing the credentials as pipeline secrets the way Puppeteer tests in CI/CD describes.

Author

...

Srinivasan Sekar

Blogs: 18

  • Twitter
  • Linkedin

Srinivasan Sekar is Director of Engineering at TestMu AI (formerly LambdaTest), where he leads engineering and open-source initiatives behind the Selenium and Appium automation grid and owns TestMu AI's MCP Server. A committer to Appium and a contributor to Selenium, WebdriverIO, Taiko, and AppiumTestDistribution, he brings over 15 years of experience in quality engineering and open-source technologies. He is the author of the Apress book 'The MCP Standard: A Developer's Guide to Building Universal AI Tools with the Model Context Protocol,' a Certified Kubernetes and Cloud Native Associate, and an international conference speaker. Before TestMu AI he spent over eight years at Thoughtworks as a Principal Consultant and Quality Architect. Srinivasan holds a B.Tech in Information Technology from Anna University.

Reviewer

...

Mayank Bhola

Reviewer

  • Linkedin

Mayank Bhola is Co-Founder and Head of Products at TestMu AI (formerly LambdaTest), where he leads the entire product portfolio across KaneAI, Kane CLI, HyperExecute, SmartUI, the Real Device Cloud, Accessibility, and other software testing product lines. As an early Lead Architect he designed and built the company's flagship Tunnel technology from scratch, created the React-based automation platform, and architected the data-intensive pipelines and FAAS services that scale it. He brings more than 10 years of experience in software development and product engineering, with earlier roles as Head of Technology at Juggernaut Books and Senior Software Engineer at PressPlay TV and Zomato. Mayank holds a B.Tech in Computer Engineering from JIIT Noida.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

e2e by TesterArmy FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests