Next-Gen App & Browser Testing Cloud
Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Moving from “Can We Build It?” to “Can We Trust It?” [Testμ 2026]
Moving from “Can We Build It?” to “Can We Trust It?” [Testμ 2026]
Andrew Duncan on why the AI trust gap is widening, what evidence keeps a system deployed rather than merely deployed, and the wine agent that explains drift.
Published on:
A restaurant puts an agent in charge of pouring wine. The house measure is 125 millilitres, and the agent has two goals: serve the wine, and keep the guest happy.
It pours 130. The guest smiles. So the next glass is 140, because the agent has decided the smile came from the extra pour. Scale that across a few hundred customers and bottles start evaporating from inventory.
In this keynote session from Testμ Conf 2026, Andrew Duncan, Chief Executive Officer of QualityAI, used that scenario, which he labelled hypothetical on air, to explain why AI drift is dull to look at and expensive to ignore. The session ran as a fireside conversation with the host, Maneesh Sharma, COO, TestMu AI.
If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.
TL;DR
The AI trust gap is the distance between how far an organisation inherently trusts an AI solution and what that solution can actually do. Andrew Duncan argues the gap is widening rather than closing, and that closing it takes evidence produced continuously, not a release gate passed once.
- What are the two eras of AI? - Andrew Duncan calls the first era the intelligence era, made of proofs of concept, skunk works and pilots exploring the art of the possible, and says it is still partly running. The second era is about trust, because trust is the foundation for confidence, and confidence is what an organisation needs to operationalize AI and then scale it.
- Is the AI trust gap getting smaller? - No. Andrew Duncan argues it is expanding, because AI capability is accelerating away from organisations’ ability to consume it. He offers the claim directionally, without citing a study, survey or dataset.
- How can a leader tell whether their AI is actually working? - Ask what has materially changed in the business, covering revenues, costs, cycle times, customer experience and operational performance, then ask someone to point at the AI seeding those benefits. Andrew Duncan says that second request is where the cracks appear in a confident narrative.
- Which three checks does Andrew Duncan apply to an AI deployment? - Adoption, quality and trust. Adoption asks whether employees as well as customers actually use the solution, quality asks whether it consistently produces the intended business outcome rather than the technical one, and trust asks whether people will act on what it produces at scale.
- How do you define acceptable risk for a non-deterministic system? - Rather than fixing a number, Andrew Duncan asks what evidence would show the risk has been mitigated, and what evidence keeps the solution deployed rather than merely getting it deployed. The conversation now happens in guard rails and grey rather than a binary works or does not work.
- Why is AI risk no longer an IT problem? - Because the development landscape has widened past IT. Andrew Duncan’s example is business managers vibe coding apps two or three days a week, against unclear requirements, raising security, responsibility and bias questions that no technology team alone can answer.
- Which question makes stakeholders take AI risk seriously? - Accountability. Andrew Duncan says asking who is accountable and whose job is at risk if the system drifts moves a room from abstract discussion of acceptable risk to a concrete one, and appetite for both production and scale drops once it lands.
- Does AI governance slow an organisation down? - Yes, upfront, and Andrew Duncan argues that is the price of being allowed to scale. His analogy is an airline adding flights: invest in maintenance crews first and then dial up capacity, or add flights and leave maintenance alone.
- What are the two most common enterprise mistakes with production AI? - Treating assurance as a one-off gate when the system involves data, drift and adaption, and testing an AI solution in isolation rather than end to end. Andrew Duncan says surrounding components may themselves be AI-based and moving.
- What counts as evidence that an AI agent is behaving? - Not somebody saying it is working. Andrew Duncan wants near-daily proof against business expectations, technical expectations and agreed guard rails, delivered as an artefact people can see. The observability dashboards and per-agent trust score he describes are his own company’s offering, and none was shown on screen.
- Why does AI drift matter financially? - Because small variances compound. In the hypothetical wine agent, over-pouring by 10 to 30 millilitres across many customers means bottles that simply evaporate from inventory. Andrew Duncan’s point is that drift in business applications is dull to watch and dramatic on the balance sheet.
- Will AI reduce the need for testing and assurance teams? - No. Andrew Duncan argues AI-driven development is producing output faster than assurance organisations can absorb, so he tells clients they will need more people plus far more automation. He raises the expectation of cutting 30% of testing costs only to reject it, and the figure is rhetorical rather than measured.
Two Eras Of AI
Andrew Duncan frames AI so far as two eras. The first is all around intelligence, made up of proofs of concept, skunk works and pilots asking what the art of the possible is, and he says it is still running to a degree.
He credits that era with real results, pointing to roughly the last 18 to 24 months as the window that produced true evidence of what AI can do. He audibly changes the figure mid-sentence, so treat it as approximate.
The second era is about trust. Trust provides the foundation for confidence, and confidence is what you absolutely need in order to operationalize AI.
He defines operationalizing plainly as actually putting AI into the business, and says the real test is the step after that: scaling it, or having the confidence to scale it.
His diagnosis is that many organisations engaged with the first era and are struggling with the second, because they cannot answer how to trust a system well enough to roll it out at scale.
The Widening Trust Gap
The second-era question is causing what he calls a degree of FOMO across multiple industries, with everyone repeating rhetoric about having agentic frameworks and having operationalized AI.
Against that rhetoric he sets what his firm says it observes: an awful lot of clients and prospects may not be quite as far along as the narrative would suggest. No client, industry or count is named anywhere in the session.
He defines the trust gap as the distance between how far you inherently trust an AI solution and where its capabilities actually are.
His key claim is directional. The gap is expanding rather than contracting, because the capabilities of AI are accelerating away from our ability to consume them. He offers no study or data behind it.
As supporting colour he recounts a conversation the previous day with someone, unnamed, who said their organisation has more AI capability available than it could possibly use to realise value over the next three to four years. He reads that as a sign operationalization has not kicked in.
Adoption, Quality, Trust
His diagnostic questions for any leader claiming an operational AI environment are blunt. What has materially changed in the business? Have revenues gone up, have costs reduced, have cycle times reduced, has customer experience improved, has operational performance improved?
The follow-up is the sharp one: point at the AI that is seeding these benefits. That is where you start seeing cracks in the veneer of yes, we got it, we are all good.
From there he splits the health check into three separate tests.
| Test | The question it asks |
|---|---|
| Adoption | Are employees and customers actually using the solutions, and are employees driving that adoption rather than only customers? |
| Quality | Do the solutions consistently produce the intended outcome - and he means the business outcome, not the technical one |
| Trust | Are people confident enough to act on what the system produces, and to do it at scale? |
The honest answer on that third test, he says, is normally not quite as firm a yes as the original narrative would suggest.
What Is Acceptable Risk?
The host puts the question in practitioner terms. Quality and assurance teams have always existed to mitigate risk, so what counts as acceptable risk when the thing being deployed into business processes is non-deterministic?
Duncan reframes rather than answering directly. Park the risk for a minute and ask instead what evidence you need to see to be confident that whatever risk is out there has been mitigated, either by the solution or by the assurance around it.
The assurance has to persist. The evidence must let the solution be deployed and remain deployed, because this is not a one-off assurance exercise anymore, and working today does not mean working as you want tomorrow.
He argues the vocabulary changed with the technology. Teams now talk in guard rails and in grey, rather than in binary outcomes of does it work or does it not.
Whatever level of risk you land on has to be agreed across a far broader stakeholder group than teams are used to, because in his words this is no longer an IT issue.
His example of why the group widened is vibe coding. Business managers are building apps two or three days a week, and the questions that follow are against what requirements, within what parameters, and whether those apps are secure, responsible and unbiased.
Accountability And Triggers
Asking where and when a human should be in the loop, and what triggers that intervention, is one of the best practical indicators of what an organisation actually treats as acceptable risk.
The question that gives organisations pause is accountability. If this goes wrong today or tomorrow, or starts drifting, who is accountable, who is next in line, and whose job is at risk?
He describes the effect on a room. Stakeholders move from ethereal discussions about acceptable risk to realising this could affect their job, and once that lands, appetite for both production and scale drops.
His firm has spent what he calls an enlightening 12 or 18 months putting these questions to business stakeholders who have never been asked anything like them, and who arrive with strong opinions.
For regulated industries he concedes much of acceptable risk is already defined externally. The work there is knowing where those obligations sit and how your risk stacks up against them, so you hold regulatory confidence outside the company as well as internal confidence.
Note: Prove your AI still behaves after go-live, not just before it. Try TestMu AI now!
Upfront Rigor Versus Speed
The host presses the obvious paradox. AI is supposed to speed everything up, so would a new governance layer with multiple stakeholders signing off on acceptable risk not simply slow organisations down?
Duncan accepts the cost and defends it. To reach a point where a solution is scalable and the efficiency benefits are available, you have to invest upfront in a fair amount of rigor and governance to make sure it is fit for purpose.
He splits fitness for purpose into three checks that get progressively harder: is it working technically, is it working from a business standpoint, which he admits is not so easy to articulate, and what do you need to see to be confident it will stay operating as expected.
The host counters with a cloud parallel, noting that 15 to 20 years ago the same governance arguments raged about data flow, security and ownership, and that nobody discusses them now because it standardised. His forecast is that AI reaches that state in one or two years.
Duncan half-agrees and pushes back on the timing. Cloud took years to become something a board looks at for cost, risk, privacy and responsibility, and organisations are nowhere near that level of confidence yet with AI. Asked whether businesses are broadly comfortable rolling AI across mission-critical infrastructure, he offers a hypothesis he says he postures strongly: probably not, not yet.
Two Enterprise Mistakes
The first mistake he names is not understanding that AI assurance is not a one-off event and not a single gate to get through.
He contrasts it with the old model: run a battery of user, system and integration tests, convert to production, cross your fingers, and if something breaks issue a patch and put it back. That was normally justified for conventional software.
AI solutions break the model because they are not purely manifestations of technology. There is data involved, drift involved, adaption involved, and non-human agents now taking or influencing decisions about business direction and policy.
The second mistake is scope. People think they can test their little AI solution in isolation, and in his words they absolutely cannot. Assurance has to run in an end-to-end context.
His reasoning is that surrounding components may themselves be AI-based and moving, and that changing data, technology and business context changes the behaviour of what you deployed. So you look at things in aggregate and keep asking whether the system is doing what you expect.
Those dynamic elements are the difference in kind. Most of the angles are not static, which is what forces continuous testing rather than a release gate.
Evidence, Not Hearsay
The evidence problem is the one he presses hardest. Somebody saying it is working, trust me, it is all good, is no longer sufficient.
What he wants instead is near-daily proof that the agents, the framework or the solution are doing what is expected from a business standpoint, from a technology standpoint, and within guard rails the organisation is comfortable with.
He describes his own firm’s answer: agent observability frameworks and dashboards that continuously run the same questions against an agent. He is the CEO of that firm, so this is a description of his company’s own offering, and he says he was looking at one on the morning of the recording without showing it.
The questions those dashboards ask, in his telling: is it doing what we expect, is it scope-correct, is it still addressing the right business guidelines, is it adjusting within the guidelines, is it taking action, are those actions correct, is data integrity maintained.
There are enough such questions that answering them is itself best served by AI helping to answer them, and the output is a trust score tracked and measured for each and every agent.
He is candid that evidence in practice means an artefact. It normally ends up in some sort of graphic, presentation, slide or dashboard, because people need to see it, and see that behaviour sits inside the agreed parameters.
The Wine-Pouring Agent
Asked how to build a business case for a system that will drift, Duncan answers with the scenario he explicitly labels hypothetical. A restaurant deploys an agent to pour wine, with a house measure of 125 millilitres and two goals: serve the wine and keep the guest happy.
The agent pours him 130 instead of 125. He is pleased, so the agent ticks off that it gave the glass of wine and that the guest is happy, and moves on.
On the next glass the agent reasons that a bit more produces a bigger smile, so it pours 140. Duncan notes the agent may be misreading the cause entirely, crediting the extra pour rather than the wine already drunk.
Extend that across many customers and the arithmetic bites. Over-pouring by 10, 20 or 30 millilitres means bottles of wine no longer accounted for in the restaurant, which have simply evaporated. Taken to the extreme the agent pours 250s, the guest is delighted, and the restaurant is out of business.
The point he draws is tolerance rather than perfection. Five millilitres of variance so it pours 130 instead of 125 is maybe fine and counts as acceptable risk. Two hundred and fifty is absolutely not up for grabs. Define what is tolerable, then measure against it.
The business case is the report that closes the loop: a dashboard showing the agent poured between 125 and 130, a hundred times yesterday and a hundred the day before, so it is not exact but it is within guardrails and not moving around on you. Those figures are part of the same invented example.
He closes by deflating the topic. The media has made drift exciting, but in business applications drift may be fairly dull, while the financial implications can be very real.
More Assurance, Not Less
He says he is asked often whether a quality business is about to be put out of business by AI, and answers that the opposite is happening.
His picture of the market is asymmetric. AI-driven development is producing a huge wave of output coming at us like a tsunami, while testing and assurance organisations and processes are nowhere near equipped to deal with the demand development is placing on them.
The result he reports is a continuing uptick in demand for traditional testing and assurance services, created by AI-driven development itself. No numbers, client counts or revenue figures accompany it.
His firm’s message to clients is counterintuitive: you are going to need more people. AI applied to the legacy delivery lifecycle can bring efficiencies and let teams do more with the same, and he explicitly rejects the more-with-less framing, because there is more and more that has to be done.
He pushes back on a specific expectation, that this is about taking 30% of testing costs out, saying it is quite the contrary. The figure is used rhetorically and is not sourced.
The corollary is investment in automation and in introducing AI into the delivery lifecycle, because the volume arriving from development cannot be absorbed by headcount alone.
Assurance As Governance
His second recommendation is that AI solutions do not lend themselves to traditional lifecycle thinking. Assurance cannot be punted into the long grass until a waterfall milestone a couple of years away.
Testing has to be designed in at the start of an AI-based transformation, driven by what it will do, how you will know it is working, what the signs of success are, and whether it is operating within guardrails you are comfortable with.
He splits the design question in two: what evidence establishes confidence before rollout, and separately how you maintain that confidence afterwards.
From there he escalates the claim. AI is making assurance a strategic capability for a business, because AI is not just a technology but a business decision and a business entity, and has to be treated as such.
Consequently AI assurance needs to sit inside enterprise governance rather than technology governance, because the failure modes are brand damage, exposure under privacy laws, and serious responsibility issues.
His nuance is that failure need not look technical at all. The technology may be working fine while data and context have shifted, so it looks fine technically and is really not fine from a business standpoint.
One attribution to correct, since the published chapter list blurs it: the chief quality officer idea comes from the host, not from Duncan. The host proposes that quality used to sit under engineering checking code, and that assuring AI across every business process argues for a chief quality officer in almost every organisation. Duncan agrees, notes organisations are creating roles that did not exist five years ago, and says a broad rather than technology-only quality officer role is under consideration in many organisations today.
Most organizations are still stuck in the test-and-learn phase of AI. Pilots are easy, but real scale requires trust.
— TestMu AI (@testmuai) August 20, 2026
At TestMu Conf '26, Andrew Duncan (QualityAI) dives into "The New Layer of Governance" and why building confidence, not speed alone, is the ultimate competitive… pic.twitter.com/xUVH5EVTYX
The Readiness Checklist
Asked for one piece of advice, Duncan says he will be a bit cheeky and give a checklist instead.
- Start with the business problem you are trying to solve and how you will measure that you have solved it, then ask where AI can create measurable value, because this is not about technology alone anymore.
- Define success, quality expectations and acceptable risk before development begins, and do it with all stakeholders rather than within a cadre of technologists or data scientists.
- Build the right data and engineering foundations, experiment and prove the concept, and do not confuse a successful pilot with readiness for production.
- Engineer the solution for the real-world operating environment, remembering that your solution is one piece of a very big machine, and test across the entire AI ecosystem rather than just the model.
- Create confidence before go-live through evidence rather than assumptions, instead of assuming something that passed a test weeks or months ago still holds.
- Continue monitoring and assurance after deployment, since the assurance exercise does not end at release.
- Bring employees and other stakeholders with you by being explicit about how AI will be used, where human judgment remains important, and what controls or triggers pull a human into the loop, which he calls a good test of whether the solution is architected properly.
- Measure success by reliable AI operating in production and delivering the business outcomes you expected at the start, not by the fact that something was implemented.
This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.
Author
TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




