Eval-First QA Agents: Testing Streaming Platforms at Fox Networks | TestMu 2026
Everyone's building AI agents for QA. Most ship demos. A few break production. Almost none have evals.
I'm building one at Fox right now. It runs against our streaming platforms, Apple TV, Roku, Fire TV, and I'm refusing to ship it until it can pass an eval suite as rigorous as the one I use to gate the test code it generates.
This talk is the field report. The architecture: a 38-tool MCP server, a 43-API helper library with a 3-tier resolution pipeline, and a Promptfoo eval suite currently holding the line at 88 assertions and 100% pass rate. The hard parts: hallucinated test steps, false-pass syndrome on visual checkpoints, brittle grounding on TV UIs that don't behave like web pages, and the recurring problem of an agent that's confident when it should be silent.
I'll walk through what actually works, what broke embarrassingly in production, and the design pattern that's emerging from the wreckage: eval-first agent development. Don't build the agent and bolt tests on afterward. Build the eval suite first. Use it to constrain what capabilities the agent is allowed to exercise. Promote new tools into the agent's hands only after the evals catch up to cover them.
If you're tired of "AI agent for QA" talks that are 80% architecture diagrams and 20% wishful thinking, this is the opposite of that, concrete tools, real failure modes, and a workflow you can adopt Monday morning.
The eval-first development loop for QA agents and why it changes the order you build things.
A practical agent architecture: MCP server design, helper libraries, and resolution pipelines for streaming-platform QA.
The five failure modes that keep biting me (including the ones I'm still hitting).
How to use Promptfoo to gate agent capabilities, not just score outputs.
A realistic path from "experimental agent" to "agent your team trusts to triage UHD VPF spikes at 2am".

TestMu Conf
Testμ(TestMu) Conference is TestMu AI’s (Formerly LambdaTest) annual flagship event, one of the world’s largest virtual software testing conferences dedicated to decoding the future of testing and development. Built by the community, for the community, it’s a space where you’re at the center, connecting, learning, and leading together. From deep-dive sessions on emerging trends in engineering, testing, and DevOps, to hands-on workshops and inspiring culture-driven talks, every experience is designed to keep you at the heart of the conversation.

From AI Assistants to AI Coworkers: How Engineering Teams Ship Faster with Enterprise Context
TestMu 2026
Keynote: Beyond Benchmarks - Evaluating Agents Against What They Are Actually Supposed to Do
TestMu 2026
Panel Discussion: Money Moves at Machine Speed - Trust, Risk, and Quality in Agentic Finance
TestMu 2026
From Load Testing to Reliability Engineering: Making Performance Testing Predict Production Behavior
TestMu 2026
Panel Discussion: Who Tests the Machines? QE Leaders on Quality in the Age of AI-Written Code
TestMu 2026
Fireside Chat: The Economics of AI Agents: How Startups Are Rethinking Value and Monetization
TestMu 2026
Panel Discussion: Mission-Critical Priorities in Quality Engineering: The Leader's Playbook
TestMu 2026