TestMu Conf 2026
Ship Faster. Test SmarterJoin Now
Ship Faster. Test SmarterJoin Now
SESSION

Build Your Own Benchmark: Why Agents Need Custom Evals (and How to Create Them)

AUG 21, 202608:00 - 08:45 AM (PT)45 MINS

Watch the recording

Watch on YouTube

Agents are becoming software’s primary users. The question is no longer which model ranks highest on public leaderboards - it is whether agents can reliably complete the real-world tasks your product promises.

Public benchmarks miss local context: your workflows, state, permissions, edge cases, and tool surfaces. Product failures are specific - bad defaults, missing docs, and friction that only appear when agents actually use your SDK, CLI, or MCP. Custom benchmarks solve this by turning real tasks into reproducible experiments you can rerun after every model, harness, or product change.

In this talk, we’ll discuss how anyone can define their own custom eval suite and benchmarks, what are the advantages of having them, and how to make them scalable and maintainable. You’ll see how to define tasksets, seed environments and agent harnesses, apply treatments (SDK/CLI/MCP/custom agents), add validators, and collect comparable evidence - completion scores, trajectories, metrics, and file changes - so you can compare, debug, and improve systems that work for agents.

Key Takeaways:

  • Takeaway

    A clear mental model for why public benchmarks fail their product (and what to do instead)

  • Takeaway

    A practical framework for turning real agent workflows into reproducible experiments

  • Takeaway

    Confidence that they can start building useful custom benchmarks without a research team

About the speaker

Haritha Sreedharan Nair:

Haritha has been building in the AI/ML space for the past 10 years. She is currently building Oqoqo, the first platform to bring research principles of evals and benchmarking to individuals, pioneering how agent interfaces will be built and evaluated in the future. Before being an entrepreneur, Haritha led enterprise products at Windsurf (acq. Cognition), building agentic coding tools for customers including Fortune 500s. She has built several sales tech products and developer tools during her tenure at Microsoft and has publications in reinforcement learning and neural architecture search.

More Sessions

Join the builders, testers, and innovators shaping the next generation of web experiences.
Testμ Conf 2026 is where they meet.

Register Now