Testing What Matters: Evaluating LLM Relevancy with DeepEval
Your LLM returns an answer — but is it actually relevant to what the user asked? Relevancy is the most deceptive failure mode in LLM applications: outputs look fluent, sound confident, and completely miss the point. This talk dives deep into measuring and testing relevancy using DeepEval, covering how to catch drift, hallucination disguised as helpfulness, and the gap between "correct" and "useful."
Key Takeaways:
The relevancy problem — why fluent ≠ relevant, and how LLMs confidently answer the wrong question.
Decomposing relevancy — answer relevancy, contextual relevancy, and faithfulness as distinct failure axes.
DeepEval's relevancy metrics — how they work under the hood, what they actually measure, and where they break down.
Building golden datasets — crafting test cases that catch relevancy drift without becoming brittle.
LLM-as-judge for relevancy — when model-graded relevancy is trustworthy and when human eval is unavoidable.
About the speaker
Monika Sharma:
Monika Sharma is a quality-driven software engineer with over a decade of experience architecting and executing comprehensive testing strategies across the software lifecycle. Her functional testing expertise spans UI, API, and human-to-agent interaction testing, complemented by deep cross-functional proficiency in performance and infrastructure testing. This breadth allows her to assess quality holistically — from the user-facing surface to the systems that support it beneath. A self-described quality enthusiast, Monika approaches testing not as a checkpoint but as a discipline, drawing on diverse methodologies to uncover risk and elevate product reliability. Her versatility across testing domains reflects a broader commitment to engineering excellence, one grounded in curiosity, rigor, and an appreciation for the craft in all its forms.
About
TestMu Conf
Testμ (TestMu) is the world’s largest virtual conference on agentic engineering and quality, built by the community, for the community. As AI reshapes how we build, test, and ship software, Testμ Conf is where you connect, grow, and lead: agentic workflows, autonomous quality, battle-tested AI playbooks, hands-on workshops, and the engineering culture driving it all.