The Trust Problem: Designing Quality Frameworks for AI-Generated Code
As developers lean on Copilot, ChatGPT, and similar tools, QE teams are inheriting a new class of defects, which is confident, syntactically perfect, and semantically wrong. This talk explores what it takes to test software you didn't write and can't fully predict.
Traditional QE frameworks were built on an assumption that has quietly stopped being true: that the person or system generating code understands the intent behind it. AI-generated code breaks this. It passes linters, compiles cleanly, and often looks more "correct" than human-written equivalents, while silently drifting from business logic, security posture, or regulatory intent. Drawing from work building governance and control frameworks for AI-driven systems in regulated environments like banking, this talk lays out a practical model for catching what traditional testing misses: how to design verification layers that check for semantic correctness and not just syntactic validity, where human review still has to sit in the loop, and how QE teams can build the muscle to test systems whose failure modes they haven't seen yet. If your team is shipping AI-assisted code today, this talk gives you a framework to start closing that trust gap.
Key Takeaways:
Syntactic correctness is no longer a proxy for semantic correctness. AI-generated code needs a distinct QE lens, one that tests for intent alignment, not just functional output, because the failure modes are different from human-introduced bugs.
Traditional test coverage metrics fall short with AI-generated code. Teams need layered verification, including automated checks plus targeted human review, focused on business logic and edge cases, not just pass/fail on unit tests.
Trust in AI-generated code must be engineered, not assumed. This means building explicit governance checkpoints, review gates, and escalation paths into the SDLC, treating AI as a contributor whose output requires the same scrutiny as an unfamiliar third-party vendor's code, not less.
About the speaker
Neelmani Verma:
Neelmani Verma is an Industry Principal at Infosys, and head of the Consulting group at Infosys Quality Engineering, a leader who has spent 22 years turning complex, high-stakes problems into competitive advantages for global organizations. Today, she is at the forefront of the generative AI revolution, embedding AI-driven innovation directly into her clients' most pressing challenges. Her ability to bridge technical depth with business impact has made her a trusted advisor for organizations navigating the next frontier of enterprise transformation.
About
TestMu Conf
Testμ (TestMu) is the world’s largest virtual conference on agentic engineering and quality, built by the community, for the community. As AI reshapes how we build, test, and ship software, Testμ Conf is where you connect, grow, and lead: agentic workflows, autonomous quality, battle-tested AI playbooks, hands-on workshops, and the engineering culture driving it all.