Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AI TestingAI

Pass@k vs Pass^k: Average Pass Rate Hides Flaky AI Agents

Pass@k asks whether any of k runs passed; pass^k asks whether all k did. See why average pass rate hides flaky AI agents and how many runs a pass^k gate needs.

Published on:

A GPT-4.1 agent that succeeded on 77.4% of its runs passed all five runs of the same task on only 53.0% of tasks. In pass^k terms, its pass^5 sat 24.4 points below its average pass rate, a gap that IBM Research's consistency study measured on AppWorld's 168 test_normal tasks.

Most benchmarks, IBM notes, report only the average. A release gate built on that number cannot tell the tasks an agent gets right every time from the ones it gets right only sometimes, which is how a flaky AI agent reaches production with a green build behind it.

Overview

Pass@k is the probability that at least one of k runs of a task succeeds; pass^k is the probability that all k succeed. Pass@k rises as k grows and suits work you can verify and retry. Pass^k falls as k grows, never exceeds the average pass rate, and is the number to gate a customer-facing AI agent on.

Pass@k, Mean@k and Pass^k Side by Side

  • pass@k: The chance that at least one of k independent runs of a task succeeds. For any task the agent can solve at all, it climbs toward 100% as k grows, which makes it the right metric when a checker such as a unit test can pick out the one passing run.
  • Mean@k: The average pass rate over k runs of each task, the figure most benchmarks report. Its expected value does not change with k, and it cannot show whether failures are spread across every task or concentrated in a few.
  • pass^k: The chance that all k runs of a task succeed, averaged across tasks. In a TestMu AI calculation, a task that passed 8 of 10 runs has an unbiased pass^3 of 46.7%, lower than the 51.2% that 0.8 cubed suggests.
  • Flaky versus regression: A flaky task passes and fails on the same unchanged scenario, while a regression passed consistently before a change and fails after it. A flaky task needs more runs; a regression needs a blocked release and a bisect. TestMu AI's Agent Assurance tells the two apart by comparing each scenario's saved snapshot and verdicts across runs.

What Is the Difference Between Pass@k and Pass^k?

Pass@k counts a task as passed when any one of its k independent runs succeeds, and pass^k counts it only when all k do. At k=1 both equal the per-run success rate. As k grows they move in opposite directions: pass@k climbs toward the share of tasks the agent ever solves, and pass^k falls toward the share of tasks it never fails.

  • pass@k - the tau-bench paper traces it to code generation, where unit tests verify each sample, and gives the community's definition: "the chance that at least one out of k i.i.d. task trials is successful".
  • pass^k - read "pass hat k" and proposed in the same paper (Yao, Shinn, Razavi and Narasimhan, June 2024) for agent tasks "requiring reliability and consistency like customer service": the chance that all k i.i.d. trials succeed, averaged across tasks.
  • Mean@k - the per-run pass rate averaged over k runs. IBM Research places it between the other two: pass^k is never above Mean@k, and Mean@k is never above pass@k.

Anthropic's guide to agent evals gives a worked example and says when to use each metric:

  • An agent with a 75% per-trial success rate passes all of 3 trials only about 42% of the time, because 0.75 cubed is about 0.42.
  • Pass@k fits tools where one success matters, and pass^k fits agents where consistency is essential.
MetricQuestion it answersAs k growsUse it for
pass@kDid at least one of k runs succeed?Rises toward the share of tasks solved at least onceWork a checker can verify and retry, such as code run against unit tests
Mean@kWhat share of all runs succeeded?Expected value stays the sameTracking capability from one agent version to the next
pass^kDid every one of k runs succeed?Falls toward the share of tasks that never failCustomer-facing agents and release gates, where a user gets one attempt

Is Pass@1 Enough to Ship an AI Agent?

No. Pass@1 is the average pass rate, and an average cannot tell an agent that fails a little on every task from one that fails completely on a few. Two agents that both average 50% show the difference (TestMu AI calculation):

  • Every task passes half its runs - pass^5 is 3.1%, because no task is reliable.
  • Half the tasks always pass and the rest always fail - pass^5 is 50.0%, because the reliable half never flips.

IBM's technical report on the consistency gap publishes the full AppWorld results, five runs on each of 168 tasks. The last column is a TestMu AI calculation: Mean@5 to the fifth power, which is what pass^5 would be if every task shared the same pass rate.

ReAct agent on AppWorldMean@5Pass^5GapMean@5 to the fifth (calculated)
GPT-4.1, baseline77.4%53.0%24.4 points27.8%
GPT-4.1, IBM guidelines, same task81.0%69.0%12.0 points34.9%
GPT-4.1, IBM guidelines, similar task79.5%66.0%13.5 points31.8%
GPT-OSS-120B, baseline33.9%10.1%23.8 points0.4%
  • Failures cluster by task - the measured pass^5 is nearly double what an even spread predicts (53.0% against 27.8%), so failures concentrate in a subset of tasks. IBM's report says the 24.4-point shortfall concentrates in about 32% of the benchmark, tasks that pass some runs and fail others.
  • Temperature 0.0 does not remove it - IBM ran the agent at temperature 0.0, so the variance was not ordinary sampling. A trajectory chains dozens of decisions, and IBM's write-up explains that a small per-step chance of flipping compounds into a large chance that some run goes differently.
  • A stronger model is not the whole fix - the much weaker GPT-OSS-120B showed a 23.8-point gap, close to GPT-4.1's 24.4. IBM's report notes the raw gap is capped by Mean@k: pass^5 divided by Mean@5 is 0.68 for GPT-4.1 against 0.30 for GPT-OSS-120B, about twice the consistency, and IBM still advises against reaching for a bigger model first.
  • The guideline rows come from IBM's own system - IBM generated those consistency guidelines and measured the result, so read the rows as evidence the gap can shrink, measured by the team that built the fix.

The same shape shows up in the OSWorld results covered in our guide to computer use agents. For the signals to track beside pass^k, such as task completion rate and trajectory similarity, see the AI agent reliability hub.

How Do You Calculate Pass^k From Repeated Runs?

Use the unbiased estimator from the tau-bench paper, computed per task and then averaged:

  • Fix the build and the scenario - every run of a task uses the same agent version, prompt, tools and fixtures, so the run itself is the only thing that varies.
  • Run each task n times and count its c successes - n must be at least k, and running more than k gives the estimator more samples to work with.
  • Score each task on its own - compute C(c,k) / C(n,k): the number of ways to pick k runs that all passed, divided by the number of ways to pick any k runs. The pass@k counterpart is 1 minus C(n-c,k) / C(n,k).
  • Average across tasks - the mean of the per-task values is the suite's pass^k.

The common shortcuts are biased:

  • Raising the suite average to the power k - understates pass^k whenever failures cluster by task, as the last column of the IBM table above shows.
  • Plugging each task's pass rate into (c/n)^k - overstates pass^k on small samples. A task with 8 passes in 10 runs gives 51.2% at k=3 and 32.8% at k=5, against the unbiased 46.7% and 22.2%.
  • Estimating pass@k as 1-(1-p)^k - the mirror-image mistake on the pass@k side, which Chen et al.'s 2021 Codex paper shows is biased when p is the observed pass rate.

The TestMu AI script below implements both estimators and prints the estimates, bounds, run counts and gate odds used in this article. The output under it is pasted unedited from a run on 2026-09-28 (Python 3.14.6); the bounds and run counts assume independent runs and 95% confidence.

from math import comb, ceil, log

def pass_hat_k(runs, k):
    """Unbiased pass^k (tau-bench): mean over tasks of C(c,k)/C(n,k).
    runs: list of (n, c) pairs, one per task, with n >= k."""
    return sum(comb(c, k) / comb(n, k) for n, c in runs) / len(runs)

def pass_at_k(runs, k):
    """Unbiased pass@k: mean over tasks of 1 - C(n-c,k)/C(n,k)."""
    return sum(1 - comb(n - c, k) / comb(n, k) for n, c in runs) / len(runs)

# One task, 8 passes in 10 runs
one = [(10, 8)]
for k in (3, 5):
    print(f"8/10 passes, k={k}: unbiased pass^k={pass_hat_k(one, k):.3f}  "
          f"plug-in (c/n)^k={(8/10)**k:.3f}  pass@k={pass_at_k(one, k):.3f}")

# Mean@5 raised to the 5th power vs the measured Pass^5 (IBM AppWorld rows)
rows = [("GPT-4.1 baseline", 0.774, 0.530),
        ("GPT-4.1 + guidelines, same task", 0.810, 0.690),
        ("GPT-4.1 + guidelines, similar task", 0.795, 0.660),
        ("GPT-OSS-120B baseline", 0.339, 0.101)]
for label, mean, measured in rows:
    print(f"{label}: Mean@5^5 = {mean**5:.1%}  measured Pass^5 = {measured:.1%}")

# Two agents that both average 50% per run
print(f"every task passes half its runs: pass^5 = {0.5**5:.1%}; "
      f"half the tasks always pass, half always fail: pass^5 = {0.5:.1%}")

# Zero failures in n runs: one-sided 95% upper bound on the per-run failure rate
for n in (3, 5, 10, 20, 30, 60):
    print(f"{n:>2} clean runs rule out failure rates above {1 - 0.05 ** (1 / n):.1%}")

# Runs needed to see at least one failure with 95% probability
for f in (0.10, 0.05):
    print(f"flip rate {f:.0%}: {ceil(log(0.05) / log(1 - f))} runs")

# Retry until green, up to 3 attempts, per-run pass 0.9
print(f"p=0.9, gate passes if any of 3 attempts passes: {1 - 0.1**3:.1%}; "
      f"all 3 must pass: {0.9**3:.1%}")
$ python pass_k.py
8/10 passes, k=3: unbiased pass^k=0.467  plug-in (c/n)^k=0.512  pass@k=1.000
8/10 passes, k=5: unbiased pass^k=0.222  plug-in (c/n)^k=0.328  pass@k=1.000
GPT-4.1 baseline: Mean@5^5 = 27.8%  measured Pass^5 = 53.0%
GPT-4.1 + guidelines, same task: Mean@5^5 = 34.9%  measured Pass^5 = 69.0%
GPT-4.1 + guidelines, similar task: Mean@5^5 = 31.8%  measured Pass^5 = 66.0%
GPT-OSS-120B baseline: Mean@5^5 = 0.4%  measured Pass^5 = 10.1%
every task passes half its runs: pass^5 = 3.1%; half the tasks always pass, half always fail: pass^5 = 50.0%
 3 clean runs rule out failure rates above 63.2%
 5 clean runs rule out failure rates above 45.1%
10 clean runs rule out failure rates above 25.9%
20 clean runs rule out failure rates above 13.9%
30 clean runs rule out failure rates above 9.5%
60 clean runs rule out failure rates above 4.9%
flip rate 10%: 29 runs
flip rate 5%: 59 runs
p=0.9, gate passes if any of 3 attempts passes: 99.9%; all 3 must pass: 72.9%

The first two lines show why pass@k cannot gate an agent: with 8 passes in 10 runs, pass@3 and pass@5 are both 100%, while pass^5 is 22.2%.

If your evals already run in Inspect, the UK AI Security Institute's framework, the estimator is built in. Its metrics documentation lists these reducers, selected with Epochs(n, reducer) or by combining --epochs n with --epochs-reducer reducer:

  • pass_k_{k} - the probability that all k epoch attempts succeed, which is pass^k.
  • pass_at_{k} - the probability of at least one correct sample given k epochs, which is pass@k.
  • at_least_{k} - 1 if at least k samples are correct, else 0, a threshold you can set below the epoch count when some flips are acceptable.

Choosing k From the Flip Rate You Cannot Ship

Pick k from the per-run failure rate you are not willing to ship. At 95% confidence, n clean runs in a row rule out only failure rates above 1 minus 0.05 to the power 1/n, so a small k proves much less than a green check suggests.

Clean runs in a rowFailure rates ruled out (95%)What a clean result still allows
3Above 63.2%A task that fails half its runs
5Above 45.1%A task that fails 4 runs in 10
10Above 25.9%A task that fails 2 runs in 10
20Above 13.9%A task that fails 1 run in 10
30Above 9.5%A task that fails 1 run in 20
60Above 4.9%A task that fails 1 run in 25
  • 29 runs for a 10% flip rate - that is what it takes to see at least one failure with 95% probability, and a 5% rate takes 59, as printed by the script above.
  • High k on a short list - that run count is why a high k belongs on a few high-consequence tasks, while the rest of the suite runs at a lower k and reports pass^k beside the mean.
  • Low k as a smoke check - IBM's guidance is that even k=3 will surface a gap you did not know you had, though it certifies nothing about a rare flip.
  • Ten runs as a floor - ten clean runs, the floor the AI agent reliability hub recommends, rule out failure rates above about 26%.
  • Intervals for tasks that do fail - these bounds cover clean runs only. For confidence intervals on a pass rate, see our guide to testing non-deterministic AI outputs.

The table assumes independent runs that vary the way production traffic does, and these setups undermine that:

  • Shared environment state - Anthropic's guide warns that when distinct trials fail because of the same limitation in the environment, such as limited CPU memory, they are not independent and the eval results become unreliable for measuring agent performance. Reset sandboxes, sessions and fixtures between runs.
  • No source of variation - tau-bench fixes each task's user prompt and database transitions and gets its variation from sampling the user and agent messages, the simulated user at temperature 1.0 and the agent at 0.0. A harness that replays identical user turns and cached tool responses measures only the model's own variation and misses the variation real users bring.

Why Is Pass@k Misleading for AI Agents?

Pass@k measures whether an agent can succeed in any of k runs, and a user who makes one request gets one run. IBM describes pass@k as optimistic, the right question when you can verify and retry, and pass^k as its pessimistic mirror image, where every attempt must succeed. In production, IBM notes, a workflow that succeeded once may fail the next time a user makes the same request.

A CI job that reruns a failed agent test until it goes green turns a pass^k gate into a pass@k gate. That reading is TestMu AI's, drawn from IBM's definitions, and the arithmetic below shows what it costs at a 90% per-run pass rate:

Gate policy (per-run pass rate 90%)What it measuresChance the gate goes green
Retry until green, up to 3 attemptspass@399.9%
One runpass@1, the average90.0%
All 3 runs must passpass^372.9%

Google described the same pattern for classic tests in its 2016 Testing Blog post "Flaky Tests at Google and How We Mitigate Them":

  • An opt-in flaky marker - a test marked flaky reported a failure only if it failed 3 times in a row, which "encourages developers to ignore flakiness".
  • Dismissed real failures - some failures that developers dismissed as flaky later proved to be legitimate failures caused by the code.

Pass@k is still the right metric where one success is enough:

  • Code checked by unit tests - the tests pick out the passing sample, so one success in k is a usable result, which is the setting pass@k came from.
  • Candidates filtered by a verifier - drafts, plans or queries that a checker ranks before anything reaches a user.
  • Capability ceilings - IBM's report describes pass@k as isolating capability from sampling noise, so a high-k pass@k tells you whether the agent can solve a task at all before you measure how often it does.

The same caution applies to the job's exit status: the Run Agent Assurance in CI/CD guide states that a process exit code of zero is not an agent-quality gate.

Is a Flaky Agent Test the Same as a Regression?

No. Google's 2016 post defines a flaky result as a test that shows both a pass and a fail with the same code, and for an agent the equivalent is both outcomes on an unchanged scenario. A regression is a task that passed consistently before a change and fails after it, and the two call for opposite responses: measure and stabilize a flaky task, block and bisect a regression.

Anthropic's guide sets the expectation that makes the split workable: regression evals should have a nearly 100% pass rate, so any decline signals that something is broken. The run history then decides the call.

What the run history showsCall itResponse
The task failed some runs before the change and still fails some after it, on an unchanged scenarioFlakyRaise k for that task, isolate its runs, re-grade a saved run to rule out the grader, and track its pass^k
The task passed every run before the change and fails after itRegressionBlock the release, rerun at the baseline k to confirm, then bisect the change: prompt, model, tool or data
The scenario, its criteria or the grader changed between the runsRedefinedRe-baseline, because the old history no longer compares
Failures cluster on one runner, one time window or one shared resourceEnvironmentFix the harness first, because those runs were not independent

A flip can start in the agent, the environment or the grader, so rule out the last two before blaming the first:

  • The agent - near-tie decisions that resolve differently from run to run, the mechanism IBM describes.
  • The environment - shared state, rate limits or resource limits that make one run depend on another.
  • The grader - re-grade one saved run several times, and if the verdict changes, the flip is in the grader rather than the agent.

Changes between agent versions that an unchanged aggregate score hides are a separate problem, covered in LLM regression testing.

What Is a Good Pass^k Score for a Production Agent?

There is no universal number, so set the bar by what one failure costs. Anthropic's guide lets capability evals with high pass rates graduate into a regression suite that runs continuously.

  • Report pass^k beside Mean@k - IBM's advice, because averages cannot distinguish a reliable agent from a lucky one. The gap between the two is the inconsistency you are shipping, and on IBM's GPT-4.1 baseline it was 24.4 points.
  • Hold the regression suite near 100% - any task there with a pass^k below 1 needs an owner before the next release, whether its run history marks it flaky, regressed, redefined or environment-bound.
  • Put intervals on small samples - Inspect's ci_wilson() gives a Wilson score interval for binary scores, which its documentation says stays well calibrated for small samples and proportions near 0 or 1.
  • Count only decided runs - a run the grader could not score is neither a pass nor a fail, so report it separately and keep a harness outage from reading as agent inconsistency.
  • Track per task - the tasks with pass^k below 1 are the work queue, and a suite-level number alone hides which ones they are.
  • A sound success check - pass^k is only as good as the check behind each run. The AI agent evaluation framework covers what to score on each run before you count how often the result repeats.

Separating Flaky Runs From Regressions With Agent Assurance

The flaky, regression and redefined rows of the decision table above depend on what each run was asked to do. Agent Assurance keeps a snapshot of every scenario definition beside each run's verdicts, and later edits do not rewrite it.

Comparing two runs sorts scenarios into newly failing, fixed, flaky and redefined, where flaky means the verdict flipped on an unchanged scenario, like a task that passes 8 of 10 runs. Two differing runs alone do not establish a flakiness rate, so count every run for pass^k. Reliability scenarios normally repeat, with a repeat count in the scenario file, because one sample does not establish consistency.

For an agent that acts, a flip can come down to which tool it called, and each run is graded on what that run did:

  • The call behind a flip - with hooks that report observed calls, each run's calls are set against the tools the agent declares, so the failing run's record shows which calls it made; hooks that report none leave those criteria Unable to Verify.
  • A not_called assertion - with observed calls from the profile's hooks, a scenario asserting issue_refund is not called fails on the one run in ten that calls it.
  • Judges told not to write - judges verify under an instruction not to change the system, since a check that called issue_refund would leave a refund behind for the next run.

Over k runs, the two approaches differ in what each pass rests on:

Across k runsAgent AssuranceTypical eval tool
Where the repeated tasks come fromFrom the agent's code, and reliability scenarios can set a repeat countThe same written or synthesized cases, scored on every run
What one run's pass rests onWhat that run changed; model judges score the reply too, but the agent's account of a refund never counts as the evidenceAn LLM judge or code check on that run's reply and trace
The tool call that flippedEach run's observed calls, checked against the declared toolsRecorded calls checked against the tool list you supply, when the trace captures them
A run nobody could checkUnable to Verify, outside the pass rate you compare across rerunsAn errored run, or a skip if you enabled skips

The eval column summarizes the category rather than one product, and some harnesses grade end state in sandboxes with a check script you write per task.

Install the Rook CLI from npm, then confirm the version:

npm install -g @testmuai/rook
rook --version

You need Node.js 22 or newer for this route; run rook inside the agent's repository. The Rook CLI install guide covers Homebrew and Windows too.

For Claude Code, the skill comes first, then a bounded request in chat:

npx @testmuai/rook-skill@latest install --agent claude-code
/rook Without rerunning either suite, use the saved snapshots and verdicts of my last two runs to separate new failures, fixed cases, changed criteria and possible flaky outcomes. Then propose up to three reliability scenarios with a repeat count for the possibly flaky ones on the staging profile, explain the target writes they could cause, then hold off on generating or running anything until I approve.
Note

Note: Every Agent Assurance run invokes your agent for real, with no rollback, so point repeated runs at staging with disposable fixtures.

Conclusion

Start with your ten highest-consequence agent tasks: run each one ten times against a single build, reset the environment between runs, and record the passes per task as your pass^k baseline. A task that misses any run goes on the flaky list and gets more runs before anyone calls it fixed, and a task that was clean at baseline and fails after a change blocks the release as a suspected regression until reruns at the baseline k confirm it. When those agents act on files, tools or APIs, the Agent Assurance results and evidence guide applies the same discipline to comparing runs: check scenario definitions, profile revisions, selected phases and coverage before calling a change a regression.

Author

...

Samyak Goyal

Blogs: 30

  • Linkedin

Samyak Goyal is a Senior Member of Technical Staff at TestMu AI engineering Kane CLI, the command-line tool that runs browser automation from the terminal, where a flow described in natural language executes in a real Chrome browser and returns pass or fail with shareable proof. He is a backend engineer with 4+ years of experience, previously an SDE at Innovaccer, where he built APIs, introduced Kafka, and cut deployment from weeks to hours. Samyak also builds multi-agent systems, skill-orchestration frameworks, and a personal copilot that indexes 200+ microservice repositories.

Reviewer

...

Saurabh Prakash

Reviewer

  • Linkedin

Saurabh Prakash is an Engineering Manager at TestMu AI (formerly LambdaTest), where he leads engineering on agentic AI development and scalable system architecture for the quality engineering platform. He has also contributed to Test at Scale, the company's open-source test intelligence platform. He brings over 9 years of experience across Node.js, Java, Spring, MVC, data structures, algorithms, and scalable system design, with earlier roles as SDE 2 at Zomato, Senior Software Engineer at LogicHub, and Software Development Engineer at Directi. Saurabh holds a B.Tech in Computer Science and Engineering from Delhi Technological University.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Pass@k and Pass^k FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests