Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- LLM Cost Tracking for Agent Evals: Gate on Cost per Success
LLM Cost Tracking for Agent Evals: Gate on Cost per Success
LLM cost tracking in agent evals: put cost per verified successful task next to pass rate, set budgets from repeated runs, and fail the build on a regression.
Published on:
Eval tools such as promptfoo, Langfuse and DeepEval can record what each model call cost. A release gate that reads only the pass rate lets a candidate build hold that number while every task it completes gets more expensive, and nothing turns red. LLM cost tracking changes a release decision only when cost sits in the gate itself, divided by the tasks the agent verifiably completed.
Every code block except the sample /rook request quotes published material (the OpenTelemetry registry, Anthropic's docs and engineering blog, or TestMu AI's Agent Assurance docs), and none of it is output from a run.
Overview
LLM (large language model) cost tracking for agent evals means reporting cost per verified successful task next to the pass rate: total agent spend in an eval run divided by the tasks that verifiably succeeded. A release gate on both measures fails a candidate worse than the baseline, beyond a tolerance, on one and no better on the other.
Five Steps From Per-Case Spend to a Release Gate
- Per-case cost tracking: Cost per eval case sums every model request the agent made on that case, including tool-loop turns, retries, sub-agent calls and failed attempts. Tag each request with its case, multiply token counts by a price table pinned to the run, and keep the raw counts so cost can be recomputed later.
- Evidence-backed passes: Cost per success divides agent spend by passes backed by observed evidence, such as a changed file or a written record, not by the agent's own report. Spend on failed and unchecked cases stays in the numerator. TestMu AI's Agent Assurance grades each acceptance criterion on what it observed, though a Pass can leave some criteria unverified.
- Judge spend: Tokens a model judge spends belong on their own line, or a stricter rubric reads as a costlier agent. Tag every judge request with a judge role, exclude those requests from agent cost, report judge spend beside agent spend for each run, and change the judge model or rubric in a separate commit.
- Percentile cost budgets: Run each eval case several times per build and compute cost per verified success for every repeat. Budget on a high percentile such as the 95th instead of the mean, so one cheap run cannot carry a build, and express the budget as a ratio to the last accepted baseline.
- Pareto release gate: Compare a candidate with the baseline on pass rate and cost per verified success together, treating a move inside the pass-rate tolerance or the cost budget as no change. A candidate that is costlier but better, or cheaper but worse while above the accuracy floor, passes only with a recorded decision.
Should AI Agent Benchmarks Report Cost Next to Accuracy?
Yes. The Holistic Agent Leaderboard paper (ICLR 2026) ran 21,730 agent rollouts across 9 models and 9 benchmarks, then compared cost with accuracy run by run.
- Pareto frontier - in only 1 of 9 benchmarks did the most costly model run sit on the cost and accuracy Pareto frontier.
- Reasoning effort - for 21 of 36 runs, higher reasoning effort did not improve accuracy.
An earlier paper from the same lead author, Kapoor et al. (2024), AI Agents That Matter, traced the problem to how agents are scored.
- Narrow focus on accuracy - agent benchmarks report accuracy without attention to other metrics, and as a result state-of-the-art agents are needlessly complex and costly.
- Retries raise the score - accuracy alone cannot identify progress, because methods such as retrying can improve it.
- HumanEval baseline - a simple warming strategy, retrying while raising the temperature, showed no significant accuracy difference from the best agent architecture, while Reflexion and LDB cost over 50% more and LATS over 50 times more.
For your own release gate:
- One axis - a CI gate that reads only the pass rate repeats that one-dimensional scoring inside your own pipeline.
- Second axis - keep the pass rate scored as LLM evaluation describes, and put cost per verified success next to it.
What Is Cost per Successful Task?
Cost per successful task is the total agent spend in an eval run divided by the number of tasks that verifiably succeeded. The Cost-of-Pass paper defines the same quantity per problem: the expected cost of one attempt divided by the success rate, which becomes infinite when the success rate is 0.
The more common units each miss part of what an agent task spends:
- Per token - a model with a lower price per token can still cost more per task if it takes more turns, retries more often, or reasons longer before it answers.
- Per request - an agent task spans many requests. Anthropic's guide to steering thinking and its cost notes that in a tool-use loop each request has its own max_tokens, so that cap does not bound the whole turn's spend.
- Per case - spend per eval case ignores whether the case passed, so a build that fails more often at the same spend looks unchanged.
A worked example shows how separate thresholds miss the change. The numbers are illustrative arithmetic on a 200-case suite, with no measured run behind them.
| Measure | Baseline | Candidate |
|---|---|---|
| Verified passes | 180 of 200 (90%) | 172 of 200 (86%) |
| Agent spend per case | 1.00 units | 1.08 units |
| Total agent spend | 200 units | 216 units |
| Cost per verified success | 1.11 units | 1.26 units |
- Pass-rate threshold - a tolerance of 5 points lets the candidate through, since 86% sits 4 points below 90%.
- Cost-per-case threshold - a ceiling of 10% above the baseline lets it through too, since spend per case rose 8%.
- Cost per verified success - rose from 1.11 to 1.26 units, 13% more for each task the agent completed (1.08 x 0.90 / 0.86 = 1.13).
Step 1: Set Up LLM Cost Tracking per Eval Case
For LLM cost tracking in an eval, cost per case is the sum of every model request the agent made while working on that case. Tag each request with the eval case ID before it leaves your harness, because one agent task fans out into requests the eval runner never sees directly:
- Tool-loop turns, where each tool result goes back to the model as a new request.
- Retries and repair attempts inside the agent.
- Sub-agent and planner calls that run under the same task.
- Turns spent on an attempt that ends in failure, which Step 2 keeps in the numerator.
Record token counts next to dollars:
- Tokens and dollars - Kapoor et al. (2024) recommend reporting input and output token counts in addition to dollar costs, so anyone can recalculate the cost at current prices.
- Pinned price table - price the tokens from a table pinned to the run, and store the table's version with the result.
- Standard attributes - the OpenTelemetry GenAI semantic conventions registry carries the counts as span attributes, all in Development status.
- No cost attribute - no attribute in that registry names a cost or a price, so cost is the token counts multiplied by your price table.
The attributes that matter for cost, with the registry's descriptions:
gen_ai.usage.input_tokens The number of tokens used in the GenAI input (prompt).
gen_ai.usage.output_tokens The number of tokens used in the GenAI response (completion).
gen_ai.usage.reasoning.output_tokens The number of output tokens used for reasoning (e.g. chain-of-thought, extended thinking).Reasoning tokens are billed whether or not you can read them. The Pricing section of Anthropic's guide to steering thinking and its cost warns that the billed output token count does not match the visible token count when the model thinks, and its documented usage example shows the split:
{
"usage": {
"input_tokens": 25,
"output_tokens": 348,
"output_tokens_details": {
"thinking_tokens": 312
}
}
}- Hidden reasoning - in that example, 312 of the 348 billed output tokens were reasoning.
- Billing total - the same section calls output_tokens the authoritative total used for billing, so sum that field per case, even when the visible text is short.
Your harness then needs those counts from every turn of every task. In Rook CLI, TestMu AI's command-line tool for Agent Assurance, the profile hook that invokes your agent can return them: the documented output of each execute turn has a usage field next to the reply and the observed tool calls:
{
"agent_reply": "Your order ships Tuesday.",
"conversation": "thread_abc123",
"usage": { "input": 1200, "output": 340 },
"calls": [
{ "name": "cancel_order", "arguments": { "id": "ORD-1" } }
],
"trace_url": "https://observability.example.com/trace/abc"
}- Source - the documented shape from the profiles, hooks, and lifecycle guide, quoted as written.
- usage - holds observed input and output token counts, and the guide says it enables token-economy scenarios.
- Observed values only - the guide tells you not to invent calls or token counts.
Summed per task, that spend is what AI agent observability watches in production, where cost per completed task works as a drift alert. In the eval, it becomes an input to the release decision.
Step 2: Count Only Verified Successes
The denominator decides whether cost per success means anything. If a pass can come from the agent's own report ("I've issued the refund"), every confident claim adds a success, and the cost of each one looks lower than it is.
Failures cost more, too, according to the General Agent Evaluation study.
- Setup - 5 agent architectures and 5 models, compared on 6 benchmarks.
- Finding - failed runs were systematically more expensive than successful ones; with the three closed-source models, failed runs averaged between 20% and 54% more steps depending on the architecture.
- Failed spend counts - spend on failed attempts belongs in the numerator, or cost per success understates what each completed task costs.
Set the counting rules before the first budget:
- Numerator - all agent spend in the run, including every failed case and every case that could not be checked.
- Denominator - passes backed by observed evidence of the effect, such as the file that changed, the record written, or the tool call made with the right arguments.
- Self-reports - the agent's summary of what it did stays in the record as context and never counts as a pass.
- Unchecked cases - a case the harness could not verify stays out of the denominator while its spend stays in the numerator, so every unchecked case raises cost per success.
Step 3: Keep Judge Spend Separate
A model judge spends tokens too, and if its calls land in the same usage totals as the agent's, a stricter rubric reads as a more expensive agent. Keep the judge's spend on its own line from the first run.
Anthropic's engineering post "Demystifying evals for AI agents" (January 9, 2026) separates how a task is scored from what it consumes. It lists model-based graders as more expensive than code but does not say how to account for their token use:
- Metrics beside the graders - the post says latency, token usage, cost per task and error rates can be tracked on a static bank of tasks.
- Separate keys - its illustrative YAML for a coding agent that must fix an authentication bypass lists graders under one key and tracked metrics, including token counts, under another.
task:
id: "fix-auth-bypass_1"
desc: "Fix authentication bypass when password field is empty and ..."
graders:
- type: deterministic_tests
required: [test_empty_pw_rejected.py, test_null_pw_rejected.py]
- type: llm_rubric
rubric: prompts/code_quality.md
- type: static_analysis
commands: [ruff, mypy, bandit]
- type: state_check
expect:
security_logs: {event_type: "auth_blocked"}
- type: tool_calls
required:
- {tool: read_file, params: {path: "src/auth/*"}}
- {tool: edit_file}
- {tool: run_tests}
tracked_metrics:
- type: transcript
metrics:
- n_turns
- n_toolcalls
- n_total_tokens
- type: latency
metrics:
- time_to_first_token
- output_tokens_per_sec
- time_to_last_token- Tag judge calls - mark every judge request with a role such as judge and the case ID it scored, and exclude those spans from the agent's cost.
- Report judge spend per run - print it beside agent spend in the run summary, so a rubric edit shows up as a judge cost change.
- Budget judges on their own - LangSmith, for one, documents a weekly LLM spend cap per evaluator and pauses the evaluator when it reaches the limit.
- Hold the judge fixed across a comparison - change the judge model or rubric in its own commit, so the agent's baseline stays comparable.
Building and validating the judge itself is covered in LLM-as-a-judge.
Step 4: Turn LLM Cost Tracking Into Budgets From Repeated Runs
Agent spend moves between runs of the same case, because the agent can take a different path or retry on one run and not the next. Kapoor et al. (2024) note that for agents used at scale, the variable cost, paid again on every run, outweighs the one-time cost of optimizing the agent by orders of magnitude, because an agent might be used millions of times. That makes the per-run figure the one to budget.
- Pin the price table and the judge, so only the agent can move the number.
- Run each case k times on the baseline and on the candidate, and record k with the result.
- Compute cost per verified success for each repeat of the suite, one figure per repeat.
- Budget on a high percentile such as P95 across the repeats instead of the mean, so one cheap run cannot carry a build.
- Express the budget as a ratio to the last accepted baseline, and move the baseline only when the Step 5 rule accepts a candidate, the band method that LLM regression testing applies to cost per scenario.
None of the papers cited here prescribes k or the percentile (the 2024 Kapoor et al. paper ran each agent five times for its own comparison), so choose both for your suite and write them next to the baseline.
Step 5: Gate the Release on Pass Rate and Cost per Success
A Pareto rule compares the candidate with the baseline on pass rate and cost per verified success at once, the view Kapoor et al. (2024) proposed so that accuracy and cost can be optimized jointly.
- Bands - a move inside your pass-rate tolerance, or inside the Step 4 budget for cost, counts as no change on that axis.
- Dominated - fail when the candidate is worse than the baseline on one axis and no better on the other.
- Costlier and better - pass only with a recorded decision that accepts the trade, then store the candidate as the new baseline.
- Cheaper and worse - fail when the pass rate falls below your accuracy floor; above the floor, pass only with a recorded decision that accepts the lower pass rate, then store the candidate as the new baseline.
- No worse on either axis - pass, and store the candidate as the new baseline when it improved on at least one.
Applied to the worked example and to the wider gate:
- Worked example - with the same bands as the two single thresholds (5 points of pass rate, 10% of cost), the drop from 90% to 86% counts as no change, while cost per verified success rose 13%, so the rule fails the candidate as dominated.
- Room on the cost side - on HotPotQA, Kapoor et al. (2024) cut variable cost by 53% with GPT-3.5 and by 41% with Llama-3-70B at similar accuracy by optimizing the two jointly.
- The accuracy half - an aggregate floor and a critical subset, both covered in AI evals, sit beside this rule.
Which Eval Tools Can Fail a Build on Cost?
Each tool below records or reads cost and can fail a build on a rule you write. The table follows each tool's own documentation as of September 28, 2026.
| Tool | Cost it records or reads | How a test can fail on cost | What the documented CI gate reads |
|---|---|---|---|
| promptfoo | The cost the provider reports for each output | The cost assertion fails a test case above a threshold; an unknown cost cannot be checked | Exit code 100 on any failed test case, or a pass rate below PROMPTFOO_PASS_RATE_THRESHOLD (default 100%) |
| Braintrust | Cost is one of the categories, with latency, errors and load, in its experiment comparison grade | A custom reporter decides what counts as a failure | bt eval exits non-zero when an eval throws an exception |
| Langfuse | Usage and cost of every LLM call | Raise RegressionError on any threshold you script | The documented example gates on exact-match accuracy, and its release-policy table lists quality policies only |
| DeepEval | Per-token cost and token-count fields on LLM spans, and an optional token_cost field on each test case | A custom metric that reads token_cost; none of its built-in metrics uses that field | Failing metrics fail the build |
- The gap - none of these documented defaults gates on cost per verified success, so the rule in this step is one you add, whichever tool runs your eval.
- Closest match - promptfoo's cost assertion comes closest, but it bounds each test case on its own, with no view of how many tasks passed.
Checking the Passes Behind Cost per Success With Agent Assurance
If 10 of the worked example's 172 candidate passes came from the agent's own report, the true cost per success is 1.33 units (216 / 162), not 1.26.
Agent Assurance grades each acceptance criterion on what it observed in the run, such as a changed file or a record that a read-only probe confirmed, and a scenario the evidence could not decide is Unable to Verify, kept out of the pass rate. A scenario can still pass with some criteria unverified, since a Pass means its verifiable criteria passed and none was observed to fail. Read coverage with the pass rate, and check that the criterion behind each success you count, such as the refund record, was verified. Its token_economy (cost) scenarios need your hook to return usage.
Runs record their profile, so you can rerun the baseline's scenario IDs under a cost-oriented model profile.
Before you divide spend by a pass, Rook CLI checks the tool calls behind it:
- Calls behind a pass - each call your hook observed, such as cancel_order in the Step 1 example, is matched to the agent's declared tool surface.
- Calls that must not happen - a not_called assertion rests on the calls your hook returns; if the profile advertises call evidence it never sends, an unobserved check could become a counted success.
- Judging without writes - judges are told to verify without changing anything, since calling issue_refund to look would create a success the agent never produced.
Against eval tools' defaults (a category, not one product):
| Aspect | Typical eval tool | Agent Assurance |
|---|---|---|
| Where the cases behind your cost figure come from | Written or synthesized, with the expected result or tool list you supply | Scenarios Rook CLI derives from the agent's code and the tools it declares |
| What makes a pass you divide spend by | A model judge or code check on the output, turns or trace | An observed effect; model judges grade the answer as well, but the agent's word that it cancelled an order never verifies the cancellation |
| Tool calls behind a counted pass | Calls compared with a tool list you write; to prove a counted pass changed a record, you script it per task | Observed calls matched to the declared tool surface when your hook returns them, plus files observed under the profile's declared paths |
| A pass the run could not check | An errored case, or a skip you opted into | Unable to Verify, kept out of the pass rate and so out of the successes you divide spend by |
Install the CLI and check the version:
npm install -g @testmuai/rook
rook --versionNode.js 22 or newer is required for the npm package. Start rook in the agent's repository, with the profile pointed at staging because the agent's writes are real; other install methods are in the Rook CLI install guide.
To run the check from Claude Code, add the rook skill:
npx @testmuai/rook-skill@latest install --agent claude-codeThen give it a bounded request:
/rook Show the saved staging profile for this agent and tell me whether its recorded capabilities include usage and calls, without invoking the target. Then propose three scenarios whose passes rest on observed effects, including one token_economy scenario. Wait for my approval before generating them, and show me the plan and the possible writes before invoking the target.Note: Agent Assurance (Rook CLI) derives scenarios from your agent's code and invokes the agent for real, so the passes you count come from what it did.
Conclusion
Start with the numerator: tag every request in your next eval run with its case ID and a role, agent or judge. With LLM cost tracking in place per case, divide agent spend by the passes backed by evidence, and print that figure on the same line as the pass rate.
If the agent you ship writes files, changes records or calls APIs, read the Agent Assurance results and evidence guide to see how a Pass is decided before you count one.
Author
Saurabh Prakash is an Engineering Manager at TestMu AI (formerly LambdaTest), where he leads engineering on agentic AI development and scalable system architecture for the quality engineering platform. He has also contributed to Test at Scale, the company's open-source test intelligence platform. He brings over 9 years of experience across Node.js, Java, Spring, MVC, data structures, algorithms, and scalable system design, with earlier roles as SDE 2 at Zomato, Senior Software Engineer at LogicHub, and Software Development Engineer at Directi. Saurabh holds a B.Tech in Computer Science and Engineering from Delhi Technological University.
Reviewer
Sandeep Yadav is a Senior Software Engineer at TestMu AI (formerly LambdaTest), where he builds the platform's test intelligence and AI-native engineering systems. He has architected autonomous GitHub Apps, vector-search code intelligence, and self-diagnosing QA workflows, and designed distributed platforms that process 2M+ daily test executions and 1B+ events, turning high-volume test, log, and code data into intelligent, self-optimizing systems. He works on embedding reasoning models into production infrastructure to power autonomous review, root-cause analysis, and analytics workflows. He brings over four years of engineering experience with deep expertise in the Elastic Stack, Apache Kafka, and Redis. Earlier he engineered a GDPR-compliant, end-to-end-encrypted secure web-chat application at Mithi. A Facebook Hackercup 2021 Round 2 qualifier and merit-scholarship recipient, Sandeep holds a B.Tech in Electrical Engineering from Delhi Technological University.
LLM Cost Tracking FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




