TestMu Conf 2026
Ship Faster. Test SmarterJoin Now
Ship Faster. Test SmarterJoin Now
SESSION

Testing What Matters: Evaluating LLM Relevancy with DeepEval

Your LLM returns an answer — but is it actually relevant to what the user asked? Relevancy is the most deceptive failure mode in LLM applications: outputs look fluent, sound confident, and completely miss the point. This talk dives deep into measuring and testing relevancy using DeepEval, covering how to catch drift, hallucination disguised as helpfulness, and the gap between "correct" and "useful."

Key Takeaways:

  • Takeaway

    The relevancy problem — why fluent ≠ relevant, and how LLMs confidently answer the wrong question.

  • Takeaway

    Decomposing relevancy — answer relevancy, contextual relevancy, and faithfulness as distinct failure axes.

  • Takeaway

    DeepEval's relevancy metrics — how they work under the hood, what they actually measure, and where they break down.

  • Takeaway

    Building golden datasets — crafting test cases that catch relevancy drift without becoming brittle.

  • Takeaway

    LLM-as-judge for relevancy — when model-graded relevancy is trustworthy and when human eval is unavoidable.

About the speaker

Monika Sharma:

Monika Sharma is a quality-driven software engineer with over a decade of experience architecting and executing comprehensive testing strategies across the software lifecycle. Her functional testing expertise spans UI, API, and human-to-agent interaction testing, complemented by deep cross-functional proficiency in performance and infrastructure testing. This breadth allows her to assess quality holistically — from the user-facing surface to the systems that support it beneath. A self-described quality enthusiast, Monika approaches testing not as a checkpoint but as a discipline, drawing on diverse methodologies to uncover risk and elevate product reliability. Her versatility across testing domains reflects a broader commitment to engineering excellence, one grounded in curiosity, rigor, and an appreciation for the craft in all its forms.

TESTMU-CONF 2026

GET YOUR FREE BOARDING PASS

I agree to TestMu AI's Privacy Policy, Conference Terms and Conditions.

About
TestMu Conf

Testμ (TestMu) is the world’s largest virtual conference on agentic engineering and quality, built by the community, for the community. As AI reshapes how we build, test, and ship software, Testμ Conf is where you connect, grow, and lead: agentic workflows, autonomous quality, battle-tested AI playbooks, hands-on workshops, and the engineering culture driving it all.

More Sessions

Join the builders, testers, and innovators shaping the next generation of web experiences.
Testμ Conf 2026 is where they meet.

Register Now