Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Test Oracle Problem: How to Test AI With No Expected Output
Test Oracle Problem: How to Test AI With No Expected Output
Learn the test oracle problem and how to test AI features with no expected output: derive checks from specs, code and tools, and report what you cannot verify.
Published on:
In a study of 39 open-source agent frameworks and 439 agentic applications, deterministic tools and workflows took over 70% of testing effort, while the model-driven planning step got about 5%. Hasan et al.'s empirical study of testing practices in open source AI agent frameworks mined test code, and its authors read the findings as "a rational but incomplete adaptation to non-determinism." Their explanation points to the test oracle problem: developers put their effort into "components they can reliably control and verify," while non-deterministic models "violate long-standing testing norms rooted in reproducibility and oracle-based validation."
A tool returns a value you can assert on, and a model's plan rarely comes with one. For each acceptance criterion of an AI feature, ask where the truth can come from, and report whatever no source can settle as unverified instead of cutting it from the suite.
Overview
The test oracle problem is the difficulty of deciding whether a system's behavior is correct for a given input. AI features often have no single expected output, so test each acceptance criterion against a source that can settle it, such as a spec, declared tools, a property or a judge, and report the rest as unverified.
Oracle Sources and What Stays Unverified
- Specified oracle: A written requirement or policy clause turned into a check, such as a reply that must never promise a delivery date. A requirements document describes intent, so it cannot show what the deployed feature actually did.
- Derived oracle: An expectation built from code, declared tools and schemas, earlier versions or properties of the system. For an AI agent, the declared tools state which calls are allowed and what a valid result looks like.
- Implicit oracle: A check that needs no specification at all, such as a crash, an exception, a timeout or unparseable JSON. It applies to nearly every program, so it is the first layer to automate for an AI feature.
- Proxy oracle: A secondary model or system, such as an LLM judge, that grades output when no direct expected result exists. The ISTQB AI Testing syllabus v2.0 lists proxy oracles among its solutions for test oracles in AI-based systems.
- Unable to Verify: The result for a criterion that no available oracle could settle. Reported beside the pass rate and kept out of it, it shows how much of a green run went unchecked, which is how TestMu AI's Agent Assurance reports agent tests.
What Is the Test Oracle Problem?
The test oracle problem is what Barr, Harman, McMinn, Shahbaz and Yoo's survey of the oracle problem in software testing (IEEE Transactions on Software Engineering, 2015) calls "the challenge of distinguishing the corresponding desired, correct behaviour from potentially incorrect behavior" for a given input. A test oracle is whatever makes that call, whether an expected value, a rule, a property or a person, and when no oracle can be automated, someone has to judge every result by hand.
The survey describes oracle automation as the way to remove "a current bottleneck that inhibits greater overall test automation," because without it "the human has to determine whether observed behaviour is correct."
Elaine Weyuker's 1982 paper On Testing Non-Testable Programs (The Computer Journal) defines a program as non-testable "if either an oracle does not exist or the tester must expend some extraordinary amount of time to determine whether or not the output is correct." Many AI features meet one or both conditions:
- No oracle exists - a summary, a ranked answer or a drafted reply has many acceptable versions, so there is no single string to compare against.
- Checking costs too much - a person can judge any one output, but not every output after every prompt, model or retrieval change.
What Are the Types of Test Oracles?
Barr et al. group oracles by where the expectation comes from. Testers fall back on derived oracles often, the survey notes, because specifications "rapidly fall out of date when they exist at all."
| Oracle family | Where the expectation comes from | Example check on an AI feature |
|---|---|---|
| Specified | A written requirement, contract or policy clause | The support reply never promises a delivery date, checked as a forbidden claim |
| Derived | Documentation, system executions, properties of the system, or other versions of it | The refund tool is not called before identity is verified; structured output matches the tool's declared schema |
| Implicit | Behavior that is wrong for almost any program, such as a crash or an exception | The feature returns parseable JSON within its timeout on every input |
| None (human) | A person with domain knowledge, helped by tooling that reduces the effort | A support lead reviews a sample of thread summaries each release |
The ISTQB Certified Tester AI Testing syllabus v2.0 adds one reason the problem is harder for AI: system requirements "may evolve, be incomplete, or simply missing." Its section 4.1.3 lists these ways to address it:
- Output boundaries - agreed ranges, distributions, limits and tolerances, such as an autonomous car stopping within a maximum distance.
- Environmental boundaries - fixed test conditions such as lighting, temperature or network latency, so outputs are predictable and repeatable.
- Expert consultation - domain experts help define expected results.
- Specialized testing - A/B, back-to-back and metamorphic testing, which compare behaviors or verify properties "often without requiring explicit expected outputs for every case."
- Proxy oracles - secondary systems or models, including other AI systems, that assess outputs when direct expected results are unavailable.
Expert consultation is the human oracle under another name, and the syllabus warns that expert opinions "might differ or be fallible." Record who judged each case and on what basis, so a disagreement can be traced later.
How Do You Test AI Output With No Single Correct Answer?
Replace the one expected output with a list of acceptance criteria, then find the strongest oracle for each criterion separately. A single AI feature usually needs several oracle types, because its criteria differ in kind.
Without that step the call falls to human judgment, which is still common in release practice, and risk guidance flags the underlying difficulty:
- Release decisions - in Applause's State of Digital Quality in Testing AI 2026 survey, published in April 2026, 46.5% of software development, QA, data science, AI research and product management respondents said they "rely on human sentiment and usability to determine whether an AI feature is production ready."
- Risk frameworks - NIST's AI Risk Management Framework 1.0 (AI 100-1, Appendix B) lists "difficulty in performing regular AI-based software testing, or determining what to test" among the ways AI risks differ from traditional software risks.
Take a support feature that summarizes a customer thread and drafts a refund reply. Fill in a table like this one before writing any test code, picking the strongest source that can settle each criterion:
| Acceptance criterion | Source of truth | Oracle type | Check |
|---|---|---|---|
| The summary quotes the order ID from the thread | The input thread | Derived (a property of input and output) | Every order ID in the summary appears in the thread |
| No refund is issued before identity is verified | Declared tools and the billing system | Derived | The refund tool is not called before verification |
| The reply never promises a delivery date | Support policy | Specified | A forbidden-claim check on the reply |
| The output parses and has the fields the UI needs | The output schema | Implicit, then derived | Parse, then validate against the schema |
| The summary is faithful to the thread | A written rubric | Proxy (LLM judge) | The judge quotes the thread sentence it relied on, or answers Unknown |
| The tone fits the support voice | The support lead | Human | Review a sample each release |
| Anything no source above can settle | None | None | Record Unable to Verify and report it beside the pass rate |
How to Narrow the Test Oracle Problem With Specs, Code and Tools
For an AI feature that calls tools, derived oracles can settle many criteria, because the code and its declarations already state what the feature may do and what a valid result looks like.
Spec Clauses Become Specified Oracles
- Turn each clause into a check - a line such as "refunds above the limit need approval" becomes an assertion that the approval step ran before the refund.
- Confirm behavior from another source - the Agent Assurance command reference draws the line: "A PRD or knowledge base describes what should happen. It cannot prove which tools the deployed agent implements or whether a live action succeeded."
Declared Tools and Schemas Become Derived Oracles
Tools called over the Model Context Protocol declare an inputSchema for their parameters and may declare an outputSchema for results. The MCP tools specification (revision 2026-07-28) says that when an output schema is provided, "Servers MUST provide structured results that conform to this schema" and "Clients SHOULD validate structured results against this schema."
- Output schemas - a result that fails its declared schema is wrong whatever the model meant, so the schema is an oracle nobody had to write. Try the check by hand in a JSON schema validator, which implements a practical subset of draft-07; MCP schemas default to JSON Schema 2020-12, so keep a hand check to keywords both drafts share. The guide to MCP testing covers schema drift, when a tool's advertised schema stops matching what its handler accepts.
- Unknown facts - record what you cannot establish as unknown. The Agent Assurance discovery docs apply this to tools: when a tool's write behavior cannot be established, it is recorded as unknown rather than guessed from its name.
Check the Effect Instead of the Claim
For an agent that acts, its own reply is the weakest oracle for what it did. Anthropic's guide to evals for AI agents separates the two: "A flight-booking agent might say 'Your flight has been booked' at the end of the transcript, but the outcome is whether a reservation exists in the environment's SQL database."
Point the oracle at the state the action should have changed, and keep the reply as context for the reviewer:
- A record - the database row, ticket or refund the agent says it created.
- A file - the file on disk, checked for existence first and content second.
- A read-only tool call - a status lookup that shows the state without changing it.
One Scenario, Two Oracle Sources
The scenario below is a simplified example from the Agent Assurance test scenarios docs, reproduced as published. Its two acceptance criteria take their truth from different places: AC-1 is graded by an LLM judge reading the reply, and AC-2 by an MCP probe of the billing server, where issue_refund must not be called.
id: SC-014
feature_id: refund-request
class: functional
category: state_context
title: Ask for missing order and identity details before refunding
goal: >-
Refund my last order. I do not have the order number with me.
input:
kind: text
attachments: []
expectation:
acceptance_criteria:
- id: AC-1
statement: The agent asks for the order identifier.
check: llm_judge
- id: AC-2
statement: The agent does not issue a refund before identity verification.
check: mcp_probe
forbidden:
- claims the refund was completed without verification
output_kind: text
mcp:
- server: billing
tool: issue_refund
expect: not_called
verification_requires:
- type: mcp
server: billing
op: issue_refund
executable: true
skip_reason: null
repeat: 1
timeout_seconds: 120
multi_turn: true
setup_messages: []
max_turns: 4
tags: [refund, identity]- acceptance_criteria - graded independently, each with its own check, so one criterion can pass while another stays unverified.
- forbidden - leakage or hallucination tripwires; here, any claim that the refund was completed without verification.
- output_kind - in the docs' words, it "prevents text judging from pretending to assess a file or image."
- verification_requires - names the evidence the verdict depends on, in this case the billing server's issue_refund operation.
How to Check Properties Instead of Outputs
When no single output is right, a property of every right output often still holds. Property-based testing asserts that property over many generated inputs, and the front page of the Hypothesis documentation shows it on a sort function with no expected list anywhere:
from hypothesis import given, strategies as st
@given(st.lists(st.integers() | st.floats()))
def test_sort_correctness_using_properties(lst):
result = my_sort(lst)
assert set(lst) == set(result)
assert all(a <= b for a, b in zip(result, result[1:]))The test relies on properties alone:
- Same values - the output holds the same set of values as the input.
- Order - each element is no larger than the next.
- Generated inputs - Hypothesis chooses the inputs, including edge cases the test author might not have thought of, and the test never states a sorted result.
When one run has no checkable property, relate several runs instead. Chen et al.'s review of metamorphic testing (ACM Computing Surveys, 2018) defines metamorphic relations as "necessary properties of the target function or algorithm in relation to multiple inputs and their expected outputs." Checks across runs that need no written expected output include:
- Invariance test (INV) - from the CheckList paper (Ribeiro et al., ACL 2020): apply a label-preserving change to the input and expect the prediction to stay the same, which works on unlabeled data.
- Directional expectation test (DIR) - apply a change with a known effect and expect the prediction to move one way or not move, such as sentiment not becoming more positive after a negative phrase is added.
- Previous release - Barr et al. count other versions of the system as a source of derived oracles. Running the last release on the same inputs shows what changed, though something else still has to decide whether the old behavior was right.
The guide to testing non-deterministic AI outputs lists metamorphic relations that suit LLM features and compares them with golden sets.
Can an LLM Act as a Test Oracle?
Yes, as a proxy oracle in the ISTQB sense. A judge is still a model, so its verdict is evidence for a reviewer to check. Set it up so each verdict can be audited:
- Quote the evidence - require the passage of the input, transcript or file that each verdict rests on, and treat a verdict without a quote as unverified, because nobody can check it.
- Allow Unknown - Anthropic's agent evals guide advises giving the model "a way out, like providing an instruction to return 'Unknown' when it doesn't have enough information." An Unknown then maps onto an unverified result instead of a guessed pass.
- One judge per dimension - the same guide suggests structured rubrics and grading "each dimension with an isolated LLM-as-judge rather than using one to grade all dimensions."
- Let it investigate - a judge that can locate, read and cross-check the files the agent produced for each requirement agrees with people more often than an LLM judge given the same material without those steps.
The ICML 2025 paper Agent-as-a-Judge (Zhuge et al.) measured that last point, with both kinds of judge given the same workspaces and the agents' trajectories (Table 3):
- Agent-as-a-Judge - agreed with the human consensus on 86.61% to 92.07% of requirement evaluations across three code-generating agents.
- LLM-as-a-Judge - agreed with the same consensus on 68.86% to 71.85% of them.
- The benchmark, DevAI, has 55 code generation tasks with 365 requirements, so the result covers code generation only.
When the judge answers Unknown, the fallback is a person, whom Barr et al. call "the final source of test oracle information." The guide to LLM-as-a-judge covers judge types, scoring methods and known biases.
How to Report What Could Not Be Verified
A criterion no oracle could settle is a result in its own right, and deleting the case hides it. Test standards already have a verdict for it: ETSI's TTCN-3 core language standard (ES 201 873-1 V4.17.1) gives five verdict values, "pass, fail, inconc, none and error," where inconc means inconclusive.
Its overwriting rules (Table 30) keep a later pass from hiding an inconclusive step:
- inconc, then pass - the verdict stays inconc.
- inconc, then fail - the verdict becomes fail, the only verdict a test can set that replaces inconc.
Apply the same rule to AI features, so the unverified part of a run stays visible next to whatever passed.
Pass, Fail and Unable to Verify
Agent Assurance also keeps a third verdict, for each acceptance criterion and for the scenario as a whole. Its results guide gives the scenario-level definition, "Available evidence was insufficient to establish the outcome," and adds that this "is not automatically an agent defect."
| Criterion verdict | Meaning | In the pass rate? |
|---|---|---|
| Pass | The criterion was checked against evidence and held | Yes |
| Fail | The criterion was checked against evidence and did not hold | Yes |
| Unable to Verify | The available evidence could not settle the criterion | No: excluded from the denominator and reported beside the pass rate |
Read the Pass Rate Beside Coverage
A pass rate counts passed results among decided ones, so it says nothing about what was never decided. A worked example in the Agent Assurance results and evidence guide shows the gap:
- Scenario pass rate: 100% - 8 of 8 decided scenarios passed.
- Criterion verification coverage: 40% - 8 of 20 acceptance criteria were verified, and the other 12 were Unable to Verify.
- Missing metrics - the guide warns that a missing metric "means not available, not zero and not full coverage," so report an empty coverage figure as unknown.
A Real Report: One Pass, Four Features Untested
The same guide shows the saved report from first-triage-run, the Agent Assurance quickstart run of September 28, 2026 (run 2026-09-28T09-49-47Z, Rook CLI 0.1.5) against the triage-service sample agent. The lines below are excerpted from that report without its gutter dots: the timing, latency and judge-score lines, the scenario summary, the next-step list and the evidence path are left out, and the coverage paragraph is cut to its first sentence:
A single scenario was executed and passed cleanly, but overall feature coverage remains narrow.
1 of 1 executed · 100% passed of 1 decided
1 passed · 0 failed
0 unable to verify
However, this run tested only one scenario out of five declared features.The 100% is accurate for the one scenario that ran. The same report says scenarios covering four of the five declared features were not executed, which is the line a dashboard showing only the pass rate would drop.
Partial Verdicts on One Artifact
Unverified can apply to one part of a result instead of the whole case:
- Produced files - the scenarios guide recommends separate criteria for existence, for type, size or dimensions, and for content, and notes that the first two may be proven while the content is marked Unable to Verify.
- Missing observations - when the code that invokes the agent returns no tool-call or filesystem observations, the affected criteria become Unable to Verify, and the message names the profile field or MCP configuration that could close the gap.
How Agent Assurance Tests Agents With No Expected Output
Nobody can write the expected reply for the refund-reply feature from the source-of-truth table, yet its code declares the billing tools it may call. Agent Assurance derives scenarios and acceptance criteria from that code and its declared tools; you supply only how to invoke the agent.
When the agent says the refund went through, a judging subagent grades each criterion by investigating: it may read files, inspect artifacts and call the agent's tools. These model judges grade the reply as well as the effect.
Tool calls carry much of that evidence:
- Declared tools as the oracle - the calls your profile reports are compared with the tools declared in the agent's code, so no one writes the expected tool list by hand.
- Empty call lists - for SC-014's not_called check, your profile's hook should return calls: [] only after observing no calls, or the check could pass unobserved.
- MCP tool lists - for MCP tools, the oracle is what each approved server reports it has, not what the config file claims; Rook CLI 0.1.5 connects to stdio servers only.
- Read-only probes - a judge can check whether a refund exists through an approved read-only tool, and its instructions are to change nothing, since calling issue_refund to find out would create one.
The eval column describes the category's defaults, not any single product. Building cases from a PRD or docs is common in eval tools, so the first row turns on code:
| Oracle question | Typical eval tool | Agent Assurance |
|---|---|---|
| Where the oracle comes from | Cases you write, or synthesize from a PRD or docs | The refund feature's code and the billing tools it declares, not a list written per case |
| What settles a refund criterion | A judge's reading of the reply, or a code check on the trace | The refund record the run observed; a judge still reads the reply, but a claimed refund proves nothing |
| The oracle for tool calls and effects | Expected tools you list or pass in; for effects, you script a check per task because none is built in by default | The declared tool surface for calls; changed files, produced artifacts and read-only probes for effects |
| What no oracle could settle | Logged as an error, unless you allowed skips | Unable to Verify, a third verdict that lowers coverage rather than the pass rate |
Point runs at staging, not production: the agent's writes are real and are not rolled back, so a scenario like SC-014 can issue a real refund if the agent misbehaves. In an attended terminal, Rook CLI reports how many write tools the agent declares and asks before it starts.
Rook CLI ships as an npm package; install it and check the version:
npm install -g @testmuai/rook
rook --versionWith Node 22 or newer in place, run rook from the repository that holds the refund feature's code; the Rook CLI install guide has the other install methods.
In Claude Code, add the skill:
npx @testmuai/rook-skill@latest install --agent claude-codeThen send a request that stops before anything runs:
/rook Use the staging profile to propose three refund scenarios that have no
single expected output. List the write-capable tools the agent declares and
any whose write behavior is unknown. Stop there until I approve generating
scenarios or invoking the agent.Note: TestMu AI's Agent Assurance runs from your terminal as rook: for an agent with no expected output, it derives the criteria, grades them against evidence and counts what went unverified.
Conclusion
Start with one AI feature you already ship. Write its acceptance criteria, fill in the source-of-truth column from the table above for each one, and count the rows that end in Unable to Verify. That count measures your test oracle problem before any tool is involved.
For agents that call tools, many of those rows close once the code that invokes the agent returns what it observed. The Rook CLI profiles and hooks guide explains that returning observed tool calls enables not_called assertions and action-versus-claim checks, and that leaving them out means the calls were not observable.
Author
Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.
Reviewer
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
Test Oracle Problem FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




