TestMu 2026 Home / Video /

Keynote: Beyond Benchmarks - Evaluating Agents Against What They Are Actually Supposed to Do | TestMu 2026

Keynote: Beyond Benchmarks - Evaluating Agents Against What They Are Actually Supposed to Do | TestMu 2026

...Playlist

...

About the talk

Agent evaluation has moved fast over the past year, bringing trace-level scoring, multi-turn and tool-use testing, LLM-as-judge calibration, and evals wired into CI rather than run once before launch. But one gap keeps showing up: the evaluations teams need for their own system usually don't exist yet. Generic metrics like helpfulness, groundedness, and toxicity are useful signals, yet a system can score well on all of them while still violating the rules that actually matter, such as issuing a refund above threshold, ignoring an approval boundary, or following an instruction hidden in a tool result.

This talk covers where agent evaluation stands today and what has genuinely changed, then focuses on the shift I think matters most: treating your written behavior specification as a first-class input to evaluation rather than background context. I'll walk through ASSERT (Adaptive Spec-driven Scoring for Evaluation and Regression Testing), an open-source framework from Microsoft's Responsible AI team that turns natural-language behavior requirements into executable evals, generating a reviewable behavior taxonomy, stratified test cases, full agent traces, and per-case verdicts with rationale and policy citations. Using a multi-agent travel-planning agent as the worked example, we'll look at where it passes, where it fails, and why an aggregate score would have hidden both.

Key Takeaways:

What is actually new in agent evaluation: trace-level scoring, multi-turn and tool-use coverage, judge calibration, and evals as a release gate.

Why generic benchmarks miss application-specific failures, and how to tell when you have outgrown them.

How to turn a requirements doc, policy, or system prompt into a running eval suite, and why specification quality determines coverage more than generation does.

Why decomposed results beat a single number.

Where spec-driven evals fall short. Vague specs produce vague scenarios, synthetic cases miss production failure modes, and judge models vary in strictness. Evals complement human review and telemetry, they don't replace them.

More from TestMu Conf 2026

Session details, speaker bio, and abstract

All 87 TestMu Conf 2026 recordings

TestMu Conf

TestMu Conf

Testμ(TestMu) Conference is TestMu AI’s (Formerly LambdaTest) annual flagship event, one of the world’s largest virtual software testing conferences dedicated to decoding the future of testing and development. Built by the community, for the community, it’s a space where you’re at the center, connecting, learning, and leading together. From deep-dive sessions on emerging trends in engineering, testing, and DevOps, to hands-on workshops and inspiring culture-driven talks, every experience is designed to keep you at the heart of the conversation.

More Videos from TestMu 2026

LT Video

Welcome to TestMu Conference 2026

TestMu 2026
LT Video

Keynote: Testing In The Era of AI - From QA to Trust Engineering

TestMu 2026
LT Video

Building Deterministic Infrastructure for Non Deterministic AI Agents

TestMu 2026
LT Video

Confidence - Correctness The Agentic Validation Loop

TestMu 2026
LT Video

Ask Me Anything: AI Agents, Skills, Career & Growth

TestMu 2026
LT Video

Context That Dreams

TestMu 2026
LT Video

From AI Assistants to AI Coworkers: How Engineering Teams Ship Faster with Enterprise Context

TestMu 2026
LT Video

Architects of Trust: Redefining Quality Leadership in an Agentic World

TestMu 2026
LT Video

When Software Becomes Everyone's Job

TestMu 2026
LT Video

Fireside Chat: AI Trust and Governance in QE - Who Actually Signs Off

TestMu 2026
LT Video

Can You Trust Your ChatBot: Techniques for Testing LLM Responses

TestMu 2026
LT Video

Flaky, Fragile, and Forgotten: The Fall of a Test Automation Project

TestMu 2026
LT Video

The Right Model for the Right Job: Cutting LLM Costs Without Cutting Quality

TestMu 2026
LT Video

System Prompt Design for Voice Agents

TestMu 2026
LT Video

Testing at the Speed of Agents: Salesforce's Evolving Approach to Agentic Testing

TestMu 2026
LT Video

Panel Discussion: The Agentic Software Factory That Enterprises Actually Need

TestMu 2026
LT Video

The Bionic Workforce: Where Humans and AI Co-exist

TestMu 2026
LT Video

Keynote: Beyond Benchmarks - Evaluating Agents Against What They Are Actually Supposed to Do

TestMu 2026
LT Video

Build Trustworthy AI Agents powered by Evals

TestMu 2026
LT Video

Real World Lessons from Scaling Test Automation at Microsoft

TestMu 2026
LT Video

Canal Mania and the Philosopher's Stone

TestMu 2026
LT Video

Why RL Environments Are All You Need for Building Agents

TestMu 2026
LT Video

Panel Discussion: QE at Enterprise Scale - What Will a QE Org Look Like in 2027?

TestMu 2026
LT Video

Working With Developers: A Tester's Survival Guide

TestMu 2026
LT Video

Testing What Matters: Evaluating LLM Relevancy with DeepEval

TestMu 2026
LT Video

Beyond Scoring: Rethinking Ranking in the LLM Era

TestMu 2026
LT Video

Backwards Scoring: Ranking Test Suites by Which Real Incidents They Would Have Caught

TestMu 2026
LT Video

Panel Discussion: The BFSI Playbook - Scaling AI with Trust and Quality

TestMu 2026
LT Video

Keynote: Is the Developer Lifecycle Dead?

TestMu 2026
LT Video

Going Past the Vibes: How to Build Real Apps

TestMu 2026
LT Video

Bringing Mobile App Validation into the Agentic Loop

TestMu 2026
LT Video

Workshop: Advanced Web Testing with Playwright and AI

TestMu 2026
LT Video

Panel Discussion: AI in Mobile QA - What Actually Works Today?

TestMu 2026
LT Video

Local Agentic Theory for Accessible Mobile Games

TestMu 2026
LT Video

Agentic Engineering: From Writing Code to Reviewing It

TestMu 2026
LT Video

Revisit Old Problems with New Eyes

TestMu 2026
LT Video

Panel Discussion: Scaling Enterprise Practice in the Agentic Era

TestMu 2026
LT Video

Workshop: The Full Agentic QA Loop for Web Applications

TestMu 2026
LT Video

The Trust Problem: Designing Quality Frameworks for AI-Generated Code

TestMu 2026
LT Video

Why Does AI Suddenly Need Forward-Deployed Engineers?

TestMu 2026
LT Video

Stop Guessing A11y: Auto-Generate Playwright Tests from Your GraphQL Schema

TestMu 2026
LT Video

Workshop: Debug with Appium MCP

TestMu 2026
LT Video

Your AI Agent Passed Every Test. Your Humans Still Rejected It. Now What?

TestMu 2026
LT Video

You Can't assertEquals an Agent: A Tester's Guide to Agentic Quality

TestMu 2026
LT Video

Panel Discussion: Money Moves at Machine Speed - Trust, Risk, and Quality in Agentic Finance

TestMu 2026
LT Video

From Load Testing to Reliability Engineering: Making Performance Testing Predict Production Behavior

TestMu 2026
LT Video

Scaling Trust, Not Automation - Rethinking Quality for the AI Era

TestMu 2026
LT Video

Eval-First QA Agents: Testing Streaming Platforms at Fox Networks

TestMu 2026
LT Video

Workshop: Engineering RemoteXPC: Building Wireless iOS Automation with Appium

TestMu 2026
LT Video

Agentic AI Changed What We Test, And How We Test It

TestMu 2026
LT Video

Cloud Assumes You Know What a Request Will Cost

TestMu 2026
LT Video

Building Once, Running Everywhere: The Future of Portable AI Applications

TestMu 2026
LT Video

The New Skill Stack for SDETs in the Agentic Era

TestMu 2026
LT Video

Developing AI Agents for Disability and Self Determination: Lessons from Project RAISE

TestMu 2026
LT Video

Panel Discussion: Who Tests the Machines? QE Leaders on Quality in the Age of AI-Written Code

TestMu 2026
LT Video

Fireside Chat: The Economics of AI Agents: How Startups Are Rethinking Value and Monetization

TestMu 2026
LT Video

Agentic AI and the Next Decade of Quality Engineering

TestMu 2026
LT Video

When You Run Out of Requirements: What Happens When You Go All-In on Agents

TestMu 2026
LT Video

Keynote: The Next Era of AI: Moving from “Can We Build It?” to “Can We Trust It?”

TestMu 2026
LT Video

Panel Discussion: Mission-Critical Priorities in Quality Engineering: The Leader's Playbook

TestMu 2026
LT Video

Panel Discussion: Agentic Engineering and Quality in Healthcare

TestMu 2026
LT Video

Scaling Quality in a Decision Intelligence Platform: The Agentic QA Playbook

TestMu 2026
LT Video

Panel Discussion: Driving the Agentic Shift in Banking Adoption Governance and Scale

TestMu 2026
LT Video

Keynote: The Self Driving Company

TestMu 2026
LT Video

AI Applications Security Puzzle

TestMu 2026
LT Video

The Last Manual Handoff: Redesigning End-to-End Testing in the Age of AI

TestMu 2026
LT Video

Panel Discussion: Reinventing the QE Practice at Global Scale in Agentic Era

TestMu 2026
LT Video

Workshop: The Full Agentic QA Loop for Mobile Applications

TestMu 2026
LT Video

The Copilot Wrote It but Who Saves It: Surviving the AI Code Tsunami

TestMu 2026
LT Video

Evidence-Based QA: Proving Quality with Agentic Test Runs and OpenSourcing How We Do It

TestMu 2026
LT Video

ML-Driven Test Intelligence at Scale - What Works What Fails and Why It Matters

TestMu 2026
LT Video

Workshop: Legacy vs Autonomous QA Arena - Surviving the AI Driven Quality Evolution

TestMu 2026
LT Video

Building a Billion-Dollar Healthcare Product at Zebra Technologies

TestMu 2026
LT Video

Productize Yourself: AI Proof Yourself

TestMu 2026
LT Video

AI Architecture Thinking: From Idea to Impact

TestMu 2026
LT Video

Keynote: Will the Real Autonomous Agent Please Stand Up

TestMu 2026
LT Video

Redefining Test Data Strategy in the Gen AI Era

TestMu 2026
LT Video

Build Your Own Benchmark: Why Agents Need Custom Evals and How to Create Them

TestMu 2026
LT Video

From Pilots to Practice: Transforming Sales with AI

TestMu 2026
LT Video

Fast, Smart, and Fragile: Rethinking Testing Leadership in the Age of AI

TestMu 2026
LT Video

Autonomous Quality Engineering: AI Systems That Generate Execute Heal Learn and Govern Quality

TestMu 2026
LT Video

Panel Discussion: Engineering Trust Quality in AI-Driven Commerce

TestMu 2026
LT Video

Discovery Found Everything: Your Agent Still Doesnt Know What Matters

TestMu 2026
LT Video

Testing the Untestable: Turning OWASP AILLM Risks into Practical QA Checks

TestMu 2026
LT Video

When Software Starts Thinking: The Now of Quality Engineering

TestMu 2026
LT Video

AI-Powered Impact-Based Testing for Faster Safer Releases

TestMu 2026
LT Video

Closing Note TestMu Conference 2026

TestMu 2026