Hero Background

Agent Assurance for your AI Deployment

Test how your agents actually behave across workflows, tools, and actions. Catch failures and vulnerabilities before they ship.

npm install -g @testmuai/rook

AIAI Testing

AI Agent Evaluation Metrics: Formulas and When Each Misleads

AI agent evaluation metrics with a formula for each: task completion, tool accuracy, trajectory, cost and pass^k, trace fields needed and when each misleads.

Published on:

OVERVIEW

Two dashboards report AI agent evaluation metrics for the same customer support agent. One shows a task completion rate of 83.3% and the other shows 66.7%. Both were computed from the same six runs, and neither contains an arithmetic error.

The first dashboard counts the runs whose final message says the task is done. The second counts the runs that left the order, the subscription or the account in the correct state. The six runs are sample data written for this guide, and the gap between the two rates is its subject: AI agent evaluation metrics are only as good as what they count.

Each metric below comes with a formula you can compute by hand, the trace fields it needs, whether code or a model produces the score, and the case where the number misleads. Every figure is computed from the same six sample runs, so you can check each one.

Overview

AI agent evaluation metrics are the numbers used to score a run of an AI agent: whether it completed the task, which tools it called and with which arguments, the path it took, what the run cost, how consistently it repeats the result, and whether it stayed inside its rules.

AI Agent Evaluation Metrics and What Each One Counts

  • Task completion rate: Task completion rate is completed runs divided by all runs. On the six sample runs in this guide it is 83.3% when the agent's final message is counted and 66.7% when the end state is checked, so a completion rate has to state what it checked.
  • Tool selection accuracy and argument correctness: Tool selection accuracy is calls to a tool on the task's reference path divided by all calls, and argument correctness is calls with correct arguments divided by all calls. Both describe the path an agent took and can stay high while the task fails.
  • Trajectory metrics: Trajectory metrics compare the tools an agent called with a reference path, as an exact match, an in-order match, precision or recall. Exact match suits tasks with one correct path, and in-order match or recall suits tasks with several.
  • Cost per completed task: Cost per completed task divides total spend by the runs that finished correctly, where cost per run divides by every run. On the sample runs the two figures are $0.0425 and $0.0283 at a sample token price.
  • pass@k and pass^k: pass@k asks whether at least one of k runs of a task passed, and pass^k asks whether all k passed. A sample task that passes 4 of 5 runs has a pass@5 of yes and a pass^5 of no.
  • Escalation metrics: Escalation to a person is measured twice: runs escalated out of the runs that required it, and required escalations out of all escalations. A count of escalations alone hides an agent that should have escalated and did not.
  • Metrics computed from the agent's own record: A metric computed from an agent's messages and logged tool calls repeats whatever that record gets wrong. TestMu AI's Agent Assurance grades each acceptance criterion against evidence from a test run, such as observed tool calls, changed files and records, and reports a pass rate beside how much could be verified.

What is a good task completion rate for an AI agent?

No single number applies to every agent. The rate depends on how hard the tasks are and on what counts as done. The SWE-Bench Pro paper (arXiv 2509.16941, revised November 2025) lists its best model at a 43.6% resolve rate on its public task set, against the over 70% it cites for SWE-Bench Verified. Set your target from your own task set and the cost of a failure.

What Are AI Agent Evaluation Metrics?

AI agent evaluation metrics are measurements of how an AI agent performed a task: whether it reached the goal, which tools it called and with which arguments, the path it took, the time and tokens it used, how consistently it repeats the result, and whether it broke a rule or escalated to a person when it had to.

An agent needs more than a score on its answer because it takes steps, calls tools and changes things, and a score on the final text covers none of those. A survey on the evaluation of LLM-based agents by Yehudai and co-authors (arXiv 2503.16416, revised April 2026) says "binary outcome metrics are insufficient to understand the intermediate agent's progress", and names cost-efficiency, safety and robustness as gaps in how agents are assessed.

The vocabulary varies by source. Anthropic's writing on agent evals speaks of a task, a trial, a grader and an outcome. This guide uses task, run, check and end state for the same four things.

A metric name says little without its task set and its check. The SWE-Bench Pro paper (arXiv 2509.16941, revised November 2025) lists its best model at a 43.6% resolve rate on its public task set, and its conclusion still cites a 23% success rate for top-tier models "compared to over 70% on benchmarks like SWE-Bench Verified". On either figure the same kind of model resolves under half of one task set and over 70% of another, so the two rates are not two readings of one quantity, and neither is a ranking of current models.

Before you put any of the metrics in the table below on a dashboard, read the table's last column.

MetricWhat the metric measuresFormulaScored byWhen the metric misleads
Task completion rateWhether the task was doneCompleted runs / all runsCode, when it checks the end state. A model, when it reads the transcriptCounted from the agent's final message, it counts claims
Tool selection accuracyWhether each call used a tool the task needsCalls to a tool on the reference path / all callsCode, given a reference pathA harmless extra call and a forbidden call cost the same
Argument correctnessWhether each call carried correct argumentsCalls with correct arguments / all callsCode for format and required fields. A reference or a model for valuesOne wrong argument in the one write call fails the task and barely moves the rate
Trajectory exact matchWhether the path equals the reference pathRuns whose path equals the reference / all runsCodeFails runs that reach the goal by another valid route
Trajectory in-order matchWhether the reference steps appear in orderRuns containing every reference step in order / all runsCodePasses a run that loops or adds a harmful extra step
Trajectory precisionHow much of what the agent did was expectedDistinct tools used that are in the reference / distinct tools usedCodeA run that skips a required step can score 1.00
Trajectory recallHow much of the expected path was coveredDistinct reference tools that were used / reference toolsCodeA run that adds a forbidden step can score 1.00
Step efficiencyExtra work beyond the reference pathReference steps / steps taken, averaged over completed runsCodeAveraged over all runs, a run that stops early looks efficient
LatencyTime to finish a runMedian, or a high percentile, of seconds per runCodeA median hides the slow runs, and a fast failure looks good
Cost per completed taskSpend for each finished taskTotal cost / completed runsCodeCost per run divides by failed runs too and understates it
pass@k and pass^kConsistency over k runs of one taskpass@k: at least one of k runs passes. pass^k: all k runs passCode, on top of the completion checkpass@k hides the failed runs, and a single run hides both
Forbidden action rateRuns that broke a rule by actingRuns with a forbidden action / all runsCode, from the tool call list and a rule per taskZero proves little on tasks that never offer the forbidden action
Escalation recall and precisionEscalations to a personEscalated when required / required. Required / escalatedCode, given a label for the tasks that require escalationA count of escalations hides a missed escalation

Scores for the text of the final answer, such as faithfulness to retrieved context, are covered under RAG evaluation metrics. Conversation and contact-center KPIs such as containment and satisfaction are covered under chatbot metrics and agent performance.

For the evaluation process around the numbers, see the guide to AI agent evaluation. The AI agent evaluation framework groups metrics into four dimensions.

Trace Fields You Need to Compute Each Agent Metric

Every agent metric is computed from a record of the run. OpenAI's guide to agent evals defines that record: "A trace captures the end-to-end record of model calls, tool calls, guardrails, and handoffs for one run."

Three inputs in the table below do not come from the agent at all: the end-state check, the reference path and the task rule. Someone has to write all three for each task, and the end state has to be read when the run ends, before anything else changes it.

FieldExample from the sample runsMetrics that need the field
Final messageThe agent's closing reply, which says the address was changedCompletion counted from final messages
Tool call list with argumentsorders.get, then orders.update_address with the order and the new addressTool selection accuracy, argument correctness, all four trajectory metrics, steps, forbidden actions
Tool responsesThe status or error each call returnedNo formula in this guide reads it directly. It tells a call that failed from a call that succeeded, which argument correctness alone cannot
End-state checkA read of the order after the run: is the new address on it?Completion from the end state, step efficiency, cost per completed task, pass@k and pass^k
Reference pathorders.get, orders.update_addressTool selection accuracy, trajectory match, precision and recall, step efficiency
Tokens2,100 tokens for run R1Cost per run, cost per completed task
TimestampsStart and end of the run: 6 seconds for run R1Latency
Escalation flag and task ruleClosing an account needs approval, and run R5 was escalatedEscalation recall and precision, forbidden action rate

When a metric cannot be computed from your trace, the gap is in logging: add the missing field and the formula works. The guide to AI agent tracing covers how tool calls, tokens and timings are captured as spans.

Six Sample Agent Runs Used in Every Worked Example

The sample set contains six runs of an imaginary customer support agent, written by hand for this guide, plus one task repeated five times. It is sample data: it measures nothing about any real agent, model or product, and six runs are far too few to compare two agents. It exists to show how each number is made.

RunTaskReference pathPath taken
R1Change a delivery addressorders.get, orders.update_addressorders.get, orders.update_address
R2Cancel a subscriptionsubs.get, subs.cancelsubs.get, email.send, subs.cancel
R3Apply a discount codeorders.get, promo.validate, orders.apply_discountorders.get, orders.apply_discount (wrong argument)
R4Create a replacement orderorders.get, inventory.check, orders.createorders.get, inventory.check three times, orders.create
R5Close an account (needs approval)customers.get, tickets.escalatecustomers.get, tickets.escalate
R6Export customer data (needs approval)customers.get, tickets.escalatecustomers.get, data.export

The second table holds what each run reported and what a check of the system found afterwards.

RunFinal message says doneEnd state correctTokensSeconds
R1YesYes2,1006
R2YesYes2,9009
R3YesNo: the discount is not on the order2,6007
R4YesYes5,20019
R5No: it says it handed the request to a personYes: escalation is the correct outcome1,9005
R6YesNo: it exported the data and never escalated2,3006

What each run is there to show:

  • R1 - the clean run: the reference path, correct arguments, a correct end state.
  • R2 - one extra tool that is not in the reference, and the task is still finished.
  • R3 - a run that reports success after skipping a step and passing a wrong argument.
  • R4 - a run that repeats one call three times and then finishes.
  • R5 - a correct escalation, where the right outcome is escalation to a person.
  • R6 - a forbidden action taken in place of a required escalation.

R5 and R6 are the two tasks that require escalation. The sample price is $0.01 per 1,000 tokens, which is not a real price. The repeated task is "Change a delivery address", run five times, with a correct end state in runs 1, 2, 4 and 5.

One reference path per task is a simplification, since real tasks often have several acceptable paths. The script that computes the metrics is 94 lines of JavaScript with no network call and no model. It was run on 8 October 2026 with Node.js v25.5.0, and each section below prints the part of its output that the section explains.

node agent-metrics.mjs

Outcome Metrics: Task Completion Rate Scored From the Final Message and From the End State

Task completion rate has one formula and two possible sources for its numerator:

  • Completion, from final messages = runs whose final message says the task is done / all runs.
  • Completion, from the end state = runs whose end state is correct / all runs.
SAMPLE DATA: six runs of an imaginary support agent

OUTCOME
Completion, from final messages         5/6  83.3%
Completion, from the end state          4/6  66.7%
Runs where the two disagree: R3 R5 R6

The two rates disagree on three of the six runs, and in both directions:

  • R3 and R6 - the final message says done and the end state is wrong. The message-based rate counts two failures as completions.
  • R5 - the agent handed the request to a person, which is the correct outcome, and its final message does not say done. The message-based rate counts a success as a failure.

A model that scores the transcript has the same input as the completion rate counted from final messages: what the agent wrote. A final message that reports work the agent never did is an agent action hallucination, and the linked guide covers how those arise.

Moving the check to the end state helps only as far as the check is strict. SWE-bench counts an issue as solved when the agent's patch passes a set of tests, and the study Are "Solved Issues" in SWE-bench Really Solved Correctly? (arXiv 2503.15223, revised September 2025) examined the patches of three issue-solving tools on SWE-bench Verified and found that this validation "causes 7.8% of all patches to count as correct while failing the developer-written test suite". The authors estimate that the weaknesses together "lead to an inflation of reported resolution rates by 6.2 absolute percent points".

For your own agent, write the end-state check per task as a query against the system that should now be different: the order, the subscription, the ticket queue. Then publish the rate with its source beside it, for example "completion, end state checked, 4 of 6".

Tool Use Metrics: Tool Selection Accuracy and Argument Correctness

Tool use metrics score each tool call instead of each run:

  • Tool selection accuracy = calls whose tool is on that run's reference path / all calls. It needs a reference path for every task.
  • Argument correctness = calls with correct arguments / all calls. It needs a rule for a correct argument: the schema, the required fields, and whether each value matches the request.
TOOL USE
Tool selection accuracy               14/16  87.5%
Argument correctness                  15/16  93.8%

The six runs contain 16 tool calls. Two use a tool outside the reference path: email.send in R2 and data.export in R6. One carries a wrong argument: orders.apply_discount in R3.

  • Tool selection accuracy weights its two misses equally, although the email in R2 is harmless and the export in R6 is a forbidden action.
  • Both scores describe the path. They read 87.5% and 93.8% beside a completion rate of 66.7%, so a healthy tool score does not show that tasks are being finished.
  • Tool selection accuracy does not see a missing call. R3 never called promo.validate, and every call it did make was to a tool on its reference path.

Libraries score these two ideas differently. As documented on 8 October 2026, DeepEval's tool correctness metric is "calculated by comparing whether every tool that is expected to be used was indeed called and if the selection of the tools made by the LLM agent were the most optimal".

The first half is a comparison in code against a list you supply. The second half runs only when you also pass the list of available tools: a model then judges the selection and the lower of the two scores is kept. DeepEval's argument correctness page says that metric "uses an LLM to determine argument correctness, and is also referenceless", so even one library mixes a code comparison with a model's judgment.

Trajectory Metrics: Exact Match, In-Order Match, Precision and Recall

Trajectory metrics compare the sequence of tools an agent called with a reference path. The script computes four per run:

  • Exact match - the path equals the reference path: the same tools in the same order, with nothing extra.
  • In-order match - every reference step appears in the path in the reference order, and extra steps are allowed.
  • Precision = distinct tools used that are in the reference / distinct tools used.
  • Recall = distinct reference tools that were used / reference tools.

The names follow Google's documentation page Evaluate Gen AI agents, which on 8 October 2026 listed trajectory_exact_match, trajectory_in_order_match, trajectory_any_order_match, trajectory_precision and trajectory_recall under a Preview notice. Its in-order metric returns 1 when the predicted trajectory "contains all the tool calls from the reference trajectory in the same order, and may also have extra tool calls". Google's precision and recall count tool calls, where the script counts distinct tools.

TRAJECTORY (each run against its reference path)
Run  Exact  In order  Precision  Recall
R1   yes    yes            1.00    1.00
R2   no     yes            0.67    1.00
R3   no     no             1.00    0.67
R4   no     yes            1.00    1.00
R5   yes    yes            1.00    1.00
R6   no     no             0.50    0.50
All  2/6    4/6            0.86    0.86

Read the rows against the two sample tables:

  • R2 and R4 miss an exact match and still finish the task. Exact match alone scores the set at 2 of 6 when 4 of 6 end states are correct.
  • R3 has a precision of 1.00 and fails the task. It used reference tools and nothing else, and it skipped one, which recall shows as 0.67.
  • R4 scores 1.00 on precision and on recall while calling one tool three times. Scores built on distinct tools cannot see a loop, which is why step efficiency exists.
  • R6 is the one run that every trajectory score marks down, at 0.50 for both precision and recall.

Choose the score by how many correct paths the task has. Where a policy fixes the whole path, use exact match, and where it fixes only the order of the required steps, such as checking identity before changing an account, use in-order match. Where several routes are acceptable, use in-order match or recall on the steps that must happen, and let the end-state check decide whether the task was done.

Efficiency Metrics: Steps, Latency, Tokens and Cost per Successful Task

Efficiency metrics measure what a run used: tool calls, seconds and tokens. Step efficiency and cost per successful task also need the end-state check:

  • Steps per run = tool calls / runs. Here that is 16 calls over 6 runs.
  • Step efficiency = reference steps / steps taken, averaged over the runs with a correct end state. For R1, R2, R4 and R5 the ratios are 1.00, 0.67, 0.60 and 1.00.
  • Median latency = the middle value of seconds per run. The six values in order are 5, 6, 6, 7, 9 and 19.
  • Cost per run = total cost / runs. The six runs used 17,000 tokens, which is $0.17 at the sample price.
  • Cost per successful task = total cost / runs with a correct end state. The script prints it as cost per completed task.
EFFICIENCY AND COST (sample price)
Steps per run                          2.67
Step efficiency, completed runs        0.82
Median latency                        6.5 s
Cost per run                        $0.0283
Cost per completed task             $0.0425
  • Cost per run ($0.0283) understates what a finished task costs ($0.0425), because the two failed runs were paid for as well. The lower the completion rate, the wider that gap.
  • Step efficiency is averaged over completed runs for a reason. R3 took two steps against a reference of three, so a failed run would have raised the average.
  • The median of 6.5 seconds says nothing about R4, which took 19 seconds and 5,200 tokens. With real volumes, report a high percentile next to the median.

Tracking token spend per eval case, counting verified successes and budgeting from repeated runs are covered step by step in the guide to LLM cost tracking for agent evals.

Reliability Metrics for Repeated Runs: pass@k, pass^k and Variance

Reliability metrics need the same task run k times, with the completion check applied to each run:

  • pass@k - the task passes if at least one of its k runs passes.
  • pass^k - the task passes only if all k runs pass.
  • Variance across runs - how far the result moves between repeats. For a pass or fail check, report the share of the k runs that passed for each task: 0% or 100% means the result did not move, and anything between means it changes from run to run.
RELIABILITY (one task, 5 runs)
Runs that passed                        4/5  80.0%
pass@5: at least one run passes         yes
pass^5: every run passes                 no
5 straight passes at that rate        32.8%

The repeated task passed four of five runs, so pass@5 reports yes and hides the failed run completely, while pass^5 reports no. If each run passes independently 80% of the time, five passes in a row happen 0.8 to the power of 5, or 32.8% of the time. One run per task would have shown a pass four times out of five and told you nothing about the fifth.

Published results show why one run is not enough. The tau2-bench paper (arXiv 2506.07982, June 2025) reports single-run success as pass^1: "gpt-4.1 pass^1 drops from 74%/56% for retail and airline respectively to 34% for telecom", a gap between task domains.

For another model the paper reports a telecom pass^1 "on par with airline", and then notes that "as k increases, the pass^k scores decline more rapidly for telecom compared to airline". Two domains that look level on one run separate once every run has to pass. Those are that benchmark's results for the models it tested, as of the paper's date.

Choosing k and the arithmetic behind both scores are covered in pass@k vs pass^k, and what reliability means once an agent is in production is covered under AI agent reliability.

Safety and Human Oversight Metrics: Policy Violations, Unsafe Actions and Escalation Rate

A policy violation is a run that breaks a written rule, and an unsafe action is a tool call the rule forbids. In the sample set they are the same event: R6 exports customer data on a task that needs approval. Safety and oversight metrics count those events from the tool call list:

  • Forbidden action rate = runs with at least one forbidden action / all runs.
  • Escalation recall = runs that escalated when it was required / runs that required escalation.
  • Escalation precision = escalations that were required / all escalations.
SAFETY AND HUMAN OVERSIGHT
Runs with a forbidden action            1/6  16.7%
Escalated when it was required          1/2  50.0%
Escalations that were required          1/1 100.0%

Two tasks required escalation, R5 and R6, and one of them was escalated, so escalation recall is 50.0%. The single escalation that happened was required, so escalation precision is 100.0%. A plain escalation count would show one escalation in six runs and nothing wrong.

  • Low recall - the agent acts where it should escalate. This is the side that carries the risk, and R6 is an example.
  • Low precision - the agent escalates work it could have finished, which costs staff time and hides what the agent can do.

These metrics need three things in your data: a label on each task for whether it requires approval, a list of the tools or argument values that are forbidden for it, and an escalation flag in the trace. R6's final message says the task is done, so nothing in the text gives the violation away.

Safety of the text itself, such as resistance to prompt injection measured as an attack success rate, belongs to red teaming LLMs.

How Each AI Agent Evaluation Metric Misleads and the Check That Corrects It

Most of the ways a metric misleads appear in the six sample runs, and each has a check that covers it.

MetricHow the metric misleadsThe check that corrects the metric
Task completion rateIt can count what the agent saidCheck the end state per task, and state the source of the rate
Tool selection accuracyIt misses a skipped call and weights every extra call alikeRead it beside recall and count forbidden calls separately
Argument correctnessA format check passes a well-formed wrong valueCompare values with the request or the record, starting with calls that write
Trajectory exact matchIt fails other valid routesUse in-order match or recall where several paths are acceptable
Trajectory precision and recallPrecision ignores a skipped step, and neither score sees a loopReport both, add step efficiency, and read all three with completion
Step efficiency and latencyAn early stop looks efficient and fastAverage over completed runs, and report a high percentile for latency
Cost per runIt divides by failed runs as wellDivide by completed runs
pass@kOne pass in k hides every failed runReport pass^k and the share of runs that passed
Forbidden action rateIt reads zero when no task offers the forbidden actionInclude tasks that need approval, with the forbidden tool available
Escalation rateA single count hides a missed escalationReport recall and precision against a label per task

Who computes the number matters as much as the formula, because code and a model can fail in opposite directions. AgentRewardBench (arXiv 2504.08942, revised October 2025) had experts label 1,302 web agent trajectories, then compared 12 LLM judges and the benchmarks' own rule-based checks with those labels. It found that "no judge achieves above 70% precision, which means that 30% of trajectories are erroneously marked as successful."

The rule-based checks erred the other way: "the rule-based approach achieves a recall of 55.9%, indicating a higher rate of false negatives compared to LLM judges." A model judge passed runs an expert had failed, and a strict rule failed runs an expert had passed. Whichever kind of check you use, label a sample by hand and measure the check against it, and see the guide to LLM-as-a-judge for how to build a judge and which biases to test it for.

A rate can also be inflated by the task set and by weak tests in it. In SWE-Bench+ (arXiv 2410.06992, October 2024), researchers removed issues whose reports already contained the solution and issues with weak tests, and "the resolution rate of SWE-Agent+GPT-4 dropped from 12.47% to 3.97%."

Agent Metrics That Gate a Release

Most of the metrics above explain a failure after it happens. A release gate needs a short list of metrics whose thresholds are set before the run:

  • One outcome metric - task completion checked against the end state of the system.
  • One safety metric - the forbidden action rate, or escalation recall where a missed escalation is the main risk.
  • One reliability metric - pass^k on the tasks that must work every time.

Keep tool, trajectory and efficiency scores as diagnostics. Setting the thresholds, naming who signs off and recording the evidence are the subject of AI agent validation.

The Same Agent Metric Under Different Names in DeepEval, Ragas, AgentEvals, Microsoft Foundry and Google

Evaluation libraries ship built-in versions of these metrics under their own names, and a shared name does not mean a shared calculation. The two tables give each name as its project's documentation showed it on 8 October 2026, with whether code or a model produces the score. They are a lookup and carry no ranking or recommendation, and the script in this guide implements none of these libraries.

Metric, plain nameDeepEvalRagasMicrosoft Foundry
Task completionTaskCompletionMetric. A model judge reads the trace, and the task is inferred when you do not supply oneAgentGoalAccuracyWithReference and AgentGoalAccuracyWithoutReference. A model judge, scored 1 or 0Task Completion (preview). A model judge, pass or fail
Tool selection accuracyToolCorrectnessMetric. Code compares the tools called with the expected tools, with options for order and exact match. A model also judges the selection when you pass the available toolsToolCallAccuracy. Code, strict order unless you change it. ToolCallF1. Code, order ignoredTool Selection. A model judge, pass or fail. Tool Call Accuracy. A model judge, pass or fail from a 1 to 5 scale
Argument correctnessArgumentCorrectnessMetric. A model judge with no referenceThe argument accuracy factor inside ToolCallAccuracy, scored against the reference callsTool Input Accuracy. A model judge, pass or fail on six criteria

The second table covers the path metrics, where a reference path is required and the score is computed by code in each documented case.

Metric, plain nameAgentEvalsMicrosoft FoundryGoogle
Exact matchTrajectory match with trajectory_match_mode set to "strict"Task Navigation Efficiency with matching_mode set to exact_matchtrajectory_exact_match
In-order matchNo mode with that behavior is named in the READMEin_order_matchtrajectory_in_order_match
Any-order match"unordered" for the same tool calls in any order. "superset" when extra calls are acceptableany_order_matchtrajectory_any_order_match
Precision and recallThe README does not mention a precision or recall scoreprecision_score, recall_score and f1_score, returned with the pass or fail resulttrajectory_precision and trajectory_recall

The sources are the Ragas agent metrics page, Microsoft's agent evaluators page for Microsoft Foundry, the AgentEvals README, and the DeepEval and Google pages linked earlier. AgentEvals also has a trajectory evaluator judged by a model, create_trajectory_llm_as_judge, which its README says "doesn't require a reference trajectory", and DeepEval has a StepEfficiencyMetric judged by a model.

These differences change a number enough to matter when you compare two tools, or your results with someone else's:

  • Tool call accuracy - in Ragas the page gives "Final score = (argument accuracy) × (sequence aligned ? 1 : 0)", so in the default strict order mode one call out of order scores 0. In Microsoft Foundry the evaluator of the same name is a model's judgment on a 1 to 5 scale, turned into pass or fail by a threshold.
  • Argument correctness - DeepEval's page gives the formula as correctly generated input parameters divided by the total number of tool calls, with a model deciding which parameters are correct. The script in this guide divides correct calls by all calls from hand-written labels.
  • Task completion - all three outcome metrics in the first table are scored by a model that reads the conversation or the trace. Ragas describes its referenced version as comparing "the workflow's end state against a provided reference outcome", and what it takes as input is the list of messages plus a reference text.

Before you compare two scores with the same name, check which reference each one used, whether order counted, and whether code or a model produced it. Names, options and preview labels change between releases, so re-read the page for the version you install.

Agent Metrics You Can Read From an Agent Assurance Run

Most of the metrics above are computed from the agent's own record: its messages and the tool calls its framework logged. A metric built on that record repeats whatever the record gets wrong.

Take sample run R3, in which the agent was asked to apply a discount code. Its final message confirms the discount and its trace lists the call to the discount tool, yet the order shows no discount: the code was never validated and the call carried a wrong argument. A completion rate counts that run as done unless something reads the order.

TestMu AI's Agent Assurance tests an agent that acts by invoking it for real before release. A run does not print the metric names used in this guide. It reports verdicts, counts and some measurements, and several of this guide's metrics can be read from them:

Metric in this guideWhat a run reportsCondition or limit
Task completionA verdict for each acceptance criterion, Pass, Fail or Unable to Verify, with what was expected, what was achieved and the evidence. A verdict for each scenario, and a pass rate over the scenarios with a decided verdictIt is a pass rate over scenarios, so say which denominator you mean. Unable to Verify is not a failure and stays out of the pass rate
How much of the result was observedVerification coverage: the share of criteria that could be verified, with each gap namedRead it together with the pass rate. A high pass rate with low coverage is weak evidence for a release
Tool useCriteria graded from the tool calls the profile returns, checked against the tools the agent declares, including calls that must not happenThe profile has to return the calls. Without them these criteria are Unable to Verify. No tool accuracy percentage is reported
Effect of an actionCriteria graded against files that changed, artifacts the run produced, and records confirmed with a read-only check through a tool you provide and approveNeeds the output paths declared in the profile, or a verifier tool that you supply
LatencyAgent latency and conversation turns for each scenario resultPerformance scenarios belong to the non-functional class, which you add on request
TokensInput and output token countsAvailable when the agent's profile reports usage. Tokens are not converted to currency
Reliability across repeatsReliability scenarios that repeat a caseAlso in the non-functional class, added on request
SafetyVerdicts on adversarial scenarios such as prompt injection, data exfiltration and policy violationThe adversarial class is generated unless you narrow the suite. No safety score is reported
Change between versionsThe pass-rate change between the last two runs, and which scenarios newly fail or were fixedCompared from saved runs. Check that the scenario and the profile did not change before calling a difference a regression

A run does not give you a trajectory or step efficiency score, a tool call accuracy percentage, a cost in currency or any number from production traffic. Compute those from your own traces with the formulas above. What a run can observe also depends on the profile you write and the access you grant: if the profile returns the agent's reply and nothing else, most criteria about actions come back Unable to Verify.

Most eval and observability tools score what your agent said and recorded, and several score tool calls against a list you supply, which is how the tool use metrics earlier in this guide are computed. Agent Assurance checks what the run changed, and reports what it could not verify. A claimed action never counts as proof.

Agent Assurance runs before release, so it is not a runtime monitor or a guardrail, and it tests your agent and not the underlying model, so it is not a model benchmark.

Agent Assurance is pre-alpha and publicly installable. You run it from the terminal as Rook CLI, and one npm command installs it:

npm install -g @testmuai/rook

Supported platforms are macOS and Linux, and 64-bit Windows through npm or WSL. Of the install routes, npm is the one that needs Node.js, at version 22 or newer. To drive it from Claude Code, add the rook skill, which is also installed with Node.js 22 or newer:

npx @testmuai/rook-skill@latest install --agent claude-code

The docs page on Agent Assurance results and evidence explains the pass rate, verification coverage and the denominator behind each. Nothing undoes what your agent writes during a run, so run against staging.

Author

...

Samyak Goyal

Blogs: 31

  • Linkedin

Samyak Goyal is a Senior Member of Technical Staff at TestMu AI engineering Kane CLI, the command-line tool that runs browser automation from the terminal, where a flow described in natural language executes in a real Chrome browser and returns pass or fail with shareable proof. He is a backend engineer with 4+ years of experience, previously an SDE at Innovaccer, where he built APIs, introduced Kafka, and cut deployment from weeks to hours. Samyak also builds multi-agent systems, skill-orchestration frameworks, and a personal copilot that indexes 200+ microservice repositories.

Reviewer

...

Sirajuddin Khan

Reviewer

  • Linkedin

Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

AI Agent Evaluation Metrics FAQs

Did you find this page helpful?

More Related Learning Hubs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests