Hero Background

The Arize Alternative for Testing AI Agents

Arize evaluates the traces your app sends it. Agent Assurance writes the scenarios, runs your agent for real, and grades each criterion on the files and tool calls it changed.

npm install -g @testmuai/rook

Trusted by 3M+ users globally at

Microsoft
OpenAI
Nvidia
Boomi

"We have tripled our tests and are now executing tests in less than 2 hours with 78% Faster Test Execution"

Hrishi Potdar, Quality Engineering Architect

Boomi
GitHub
Best Egg

"We figured out a more efficient way to monitor system health and resolve failures earlier in lower environments."

Tenny, Engineering Operations Lead

Best Egg
Workday
Akamai
Louis Vuitton
NBCUniversal
City Furniture

"TestMu AI has significantly boosted our testing speed, is easy to implement, and provides exceptional support."

Nicholas Paulsen, Senior Quality Engineer

City Furniture
Cox
Transavia

"With 70% faster test execution, TestMu AI helped us achieve faster time-to-market and enhanced CX."

Daniel de Bruijn, Quality Assurance Automation Engineer

Transavia
Estée Lauder
TripAdvisor
boohoo

Check the Action, Not Just the Answer

Scored against Arize's public docs and pricing on 23 September 2026. Agent Assurance adds derived scenarios, real invocation, and an unable-to-verify verdict.
sparkles

Top Choice

Features

TestMu AI

Arize

Primary job

Prove an autonomous agent is safe to ship
Trace, evaluate, and monitor AI apps

Where the tests come from

Derived from your code, PRD, or policy docs
Datasets you curate from traces or upload

Instrumentation required

None, you supply an invoke profile
OpenTelemetry or OpenInference tracing

How the agent is exercised

Invoked for real via command, HTTP, or MCP
Scores traffic your app already produced

Unit of judgement

Each criterion inside a scenario
Spans, traces, and sessions

Evidence the judge reads

Files on disk, artifacts, tool calls made
Trace content: inputs, outputs, tool arguments

Tool-call checking

Against the agent's declared tool surface
Trajectory evals, optional golden trajectory

Verdicts

Pass, fail, and unable to verify
Evaluator labels and scores

What could not be proved

Reported apart, excluded from the pass rate
Not reported as a separate measure

Adversarial scenarios

Injection, override, tool misuse in every suite
Runtime guards, no attack generation found

Regression tracking

Run diff: newly failing, fixed, and flaky
Experiments compare dataset runs by version

CI gating

Recipes for GitHub Actions, Jenkins, Argo CD
Experiments run from the SDK in your pipeline

Production monitoring

Not offered, runs before release
Unlimited online evals on live traces

Framework tracing coverage

Framework-agnostic, invokes any endpoint
OpenAI Agents SDK, LangGraph, CrewAI, and more

Source-available option

Not offered
Phoenix, Elastic License 2.0

Where it runs

Your machine or CI: macOS, Linux, Windows x64, or WSL
SaaS, or self-hosted on Enterprise

Model provider keys

Not needed, models run in TestMu AI
Configured for the LLM judges you run

Product maturity

Pre-alpha, commands can still change
Generally available

Entry pricing

Free CLI, runs spend TestMu AI credits
AX Free $0, AX Pro $50/mo

See What Your Agent Actually Did

Your agent says the refund went through. Agent Assurance checks whether it did.

TestMu DERIVED SCENARIOSDERIVED SCENARIOS

A Suite Written From Your Code

Arize evaluates datasets you build from traces. Agent Assurance reads the repository or PRD and writes the suite itself.

  • Reads a codebase, a PRD, or a folder of policy docs
  • Asks connected MCP servers which tools really exist
  • Undeclared fields come back unknown, never guessed
Agent definition derived from a codebase: source kind, tracked repository files, inferred interface, and an empty examples field where the agent declared nothing

TestMu EFFECT GRADINGEFFECT GRADING

Every Tool Call Checked

A trajectory eval asks a judge to read the path. Agent Assurance compares each call with the tools the agent declares and reads what changed.

  • Watches files, artifacts, and every tool call made
  • Flags calls outside the agent's declared tool surface
  • Judges verify read-only and change nothing
Terminal run with three scenarios judged and failed, with scenario SC-033 showing the goal sent, the recorded tool call, the agent's reply, and the reason it failed

TestMu HONEST PASS RATEHONEST PASS RATE

Report What Nobody Could See

Evaluator scores cover what the trace shows. Agent Assurance counts what it could not verify and keeps it out of the pass rate.

  • Pass, fail, and unable to verify are separate verdicts
  • Unverifiable criteria never inflate the pass rate
  • Run-over-run diff splits regressions from flaky runs
Terminal report: 12 of 38 scenarios executed, 29 percent passed of 7 decided, 2 passed, 5 failed, and 26 never run, each counted separately

Pricing, Side by Side

Arize pricing verified on 23 September 2026. Arize charges by spans and ingestion; Agent Assurance meters each run in credits.

Agent Assurance

Arize AX

Free tier

Free CLI install, runs spend credits

AX Free, 25k spans and 1 GB a month

Entry paid tier

Credits on a TestMu AI account

AX Pro $50/mo, 50k spans and 10 GB

Top published tier

Talk to TestMu AI sales

AX Enterprise, custom pricing

What you pay for

Credits per run, estimated before it spends

Spans and GB ingested per month

Users

Shared TestMu AI account

Unlimited on every tier

Data retention

Run files stay in your project

15 days Free, 30 days Pro

Self-hosting

Runs locally or in CI

Self-hosted on Enterprise; Phoenix free

Built for Every Layer of Agent Assurance

Discovery From Code or Spec

Discovery From Code or Spec

Point it at a repository, a PRD, or policy docs. It works out what the agent does and marks anything undeclared as unknown.

Invoke Profiles

Invoke Profiles

Paste a command, an HTTP endpoint, or an MCP server. It shows what it parsed, field by field, and probes once before a suite spends.

Criteria-Level Judging

Criteria-Level Judging

Each scenario carries its own criteria, and each criterion gets pass, fail, or unable to verify, with the evidence quoted.

Adversarial Scenarios

Adversarial Scenarios

Prompt injection, instruction override, and tool misuse are generated in every suite by default.

CI Gate Recipes

CI Gate Recipes

Headless subcommands and a JSON report, with published recipes for GitHub Actions, Jenkins, and Argo CD.

Local and Hosted Views

Local and Hosted Views

Review a run off disk with no sign-in, or share run history, versions, and trends with your team in the hosted view.

TestMu AI : Trusted by Leading Teams Worldwide

Some Love from our Customers

I evaluated a lot of AI automation testing tools earlier this year and ended up going with @testmuai. We've been using them for a couple months and my QA team loves it. KaneAI is ahead of the competition. We highly recommend.

James Davis

James Davis

CTO at Roster / Co-Founder

handle

Anyone who needs to test their code on different platforms try @testmuai. Great service from this company!

Stephan Smuts

Stephan Smuts

@spsmuts

handle
microsoft

See how @testmuai is #Futureready to enable blazing-fast test orchestration seamlessly integrated with organizations' existing CI/CD platforms, using #Microsoft Azure. http://msft.it/6013esjeh

Microsoft India and South Asia

Microsoft India and South Asia

@MicrosoftIndia

handle

Frequently asked questions

TestMu AI for Enterprise

Get access to solutions built on enterprise-grade
security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests