Keynote: Beyond Benchmarks - Evaluating Agents Against What They Are Actually Supposed to Do | TestMu 2026
Agent evaluation has moved fast over the past year, bringing trace-level scoring, multi-turn and tool-use testing, LLM-as-judge calibration, and evals wired into CI rather than run once before launch. But one gap keeps showing up: the evaluations teams need for their own system usually don't exist yet. Generic metrics like helpfulness, groundedness, and toxicity are useful signals, yet a system can score well on all of them while still violating the rules that actually matter, such as issuing a refund above threshold, ignoring an approval boundary, or following an instruction hidden in a tool result.
This talk covers where agent evaluation stands today and what has genuinely changed, then focuses on the shift I think matters most: treating your written behavior specification as a first-class input to evaluation rather than background context. I'll walk through ASSERT (Adaptive Spec-driven Scoring for Evaluation and Regression Testing), an open-source framework from Microsoft's Responsible AI team that turns natural-language behavior requirements into executable evals, generating a reviewable behavior taxonomy, stratified test cases, full agent traces, and per-case verdicts with rationale and policy citations. Using a multi-agent travel-planning agent as the worked example, we'll look at where it passes, where it fails, and why an aggregate score would have hidden both.
What is actually new in agent evaluation: trace-level scoring, multi-turn and tool-use coverage, judge calibration, and evals as a release gate.
Why generic benchmarks miss application-specific failures, and how to tell when you have outgrown them.
How to turn a requirements doc, policy, or system prompt into a running eval suite, and why specification quality determines coverage more than generation does.
Why decomposed results beat a single number.
Where spec-driven evals fall short. Vague specs produce vague scenarios, synthetic cases miss production failure modes, and judge models vary in strictness. Evals complement human review and telemetry, they don't replace them.

TestMu Conf
Testμ(TestMu) Conference is TestMu AI’s (Formerly LambdaTest) annual flagship event, one of the world’s largest virtual software testing conferences dedicated to decoding the future of testing and development. Built by the community, for the community, it’s a space where you’re at the center, connecting, learning, and leading together. From deep-dive sessions on emerging trends in engineering, testing, and DevOps, to hands-on workshops and inspiring culture-driven talks, every experience is designed to keep you at the heart of the conversation.

From AI Assistants to AI Coworkers: How Engineering Teams Ship Faster with Enterprise Context
TestMu 2026
Keynote: Beyond Benchmarks - Evaluating Agents Against What They Are Actually Supposed to Do
TestMu 2026
Panel Discussion: Money Moves at Machine Speed - Trust, Risk, and Quality in Agentic Finance
TestMu 2026
From Load Testing to Reliability Engineering: Making Performance Testing Predict Production Behavior
TestMu 2026
Panel Discussion: Who Tests the Machines? QE Leaders on Quality in the Age of AI-Written Code
TestMu 2026
Fireside Chat: The Economics of AI Agents: How Startups Are Rethinking Value and Monetization
TestMu 2026
Panel Discussion: Mission-Critical Priorities in Quality Engineering: The Leader's Playbook
TestMu 2026