Hero Background

Next-Gen App & Browser Testing Cloud

Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

Next-Gen App & Browser Testing Cloud
Testμ

The Human Quality Layer for AI Agents [Testμ 2026]

Justin Roy of Microsoft on the five questions that decide whether people can work with an agent: can I see it, understand it, influence it, own it, trust it.

Published on:

A readiness slide in this session reports that only about a third of people understood how agents function, while 13% said agents were deeply integrated into their work and 23% did not know whether their organisation had implemented AI at all.

No source was named for those numbers on air, so treat them as the framing rather than the evidence. The behaviour they point at is the substance.

At Testμ Conf 2026, Justin Roy, Senior Security Assurance Engineer at Microsoft, argues that an agent can pass every technical evaluation and still be rejected by the people meant to use it. He presented the session; Jill Hosmer-Jolley, Faculty at California State University, is credited as co-presenter but was affected by technical problems for much of it, so the substance here is his.

Youtube thumbnail

If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.

TL;DR

The human quality layer is the set of organisational and behavioural conditions that decide whether people can actually work with an agent: visibility, understanding, influence, ownership and calibrated trust. It sits alongside model, security and reliability evaluations rather than replacing them, and it turns adoption into a quality outcome you can design, instrument, test and gate against rather than a change-management score.

  • Does passing every technical evaluation mean an agent is ready to deploy? - No. Justin Roy argues a technically sound agent still fails as a deployed system if people cannot inspect its work, challenge its decisions, control its actions or carry responsibility with real authority.
  • What are the five questions of the human quality layer? - Can I see it? Can I understand it? Can I influence it? Can I own it? Can I trust it appropriately? The wording stays plain so someone can ask them during real work and a cross-functional team can discuss quality without a specialist vocabulary.
  • What does quiet rejection of an AI agent look like? - Quiet rejection is not a complaint. Justin Roy describes people keeping the old spreadsheet, rechecking every field, asking a colleague to verify the answer, taking screenshots for a private audit trail, and using the agent only where consequences are low.
  • Does quiet rejection show up on an adoption dashboard? - No. A signed-in user, a completed run and an accepted output all read as adoption, while the workflow runs the old process and the new one side by side and gets slower.
  • Which signals reveal quiet rejection? - Verification frequency, overrides, abandoned runs, duplicate work, time to detect an error and time to recover, plus whether people can explain an outcome and say what the agent did on their behalf.
  • Can trust in an agent be announced at launch? - No. Justin Roy places trust last of the five questions because it cannot be inferred from activation numbers; it grows through repeated experience in which visibility, understanding, influence and ownership all hold.
  • What belongs in a human-readable agent run record? - Five things: current state, evidence, action history, decision boundaries and uncertainty. Gaps count as evidence too, including a failed dependency query and how fresh each source was.
  • Is more explanation the same as more understanding? - No. A long answer can sound confident and be impossible to verify. Justin Roy separates the conclusion, the supporting evidence, the uncertainty and the next verification step, and prefers reproducible artifacts over claims about invisible reasoning.
  • Should every reversal be called undo? - No. A configuration change may be reversible while a message already sent to customers is not, and Justin Roy warns against using the same comforting word for both.
  • Which accountability roles should an agentic workflow name? - The recommender, the approver, the executor, the monitor and the escalation owner, even when one person holds several. Awaiting human approval hides the real question: which human, in which role, with what authority.
  • Is a successful demo enough to raise an agent’s autonomy? - No. Autonomy should rise only when evidence across relevant conditions supports it, ambiguity and partial failure included, and not because a first incident ended well or a competitor uses the word autonomous.
  • How can a team start testing the human quality layer? - Pick one consequential workflow, bring in the people who carry the consequences, and rather than asking whether they trust the agent, ask them to show how they work and how they verify.

Skepticism as a Request

He opens with the person who does not trust the agent yet, and reframes them. Someone responsible for a customer, a decision or a risk is not resisting change; they are making a reasonable request.

Show me the evidence. Help me understand what happened. Give me a safe way to disagree. Tell me who owns the result when something goes wrong.

The easy answers, more training and more communication and more time, are sometimes right. Stopping there misses what the behaviour is telling you about the system.

Engineers already test accuracy, grounding, reliability, security, privacy and safety, and ask whether the agent completes the task, calls the right tools, stays inside policy and recovers from technical failure. Those tests are essential and they do not reach the whole production boundary.

The Five Questions

The framework is deliberately plain, so an accountable person can ask it during real work. Each question also has an engineering translation, which is what stops it being an aspiration.

The questionWhat the person needsEngineering translation
Can I see it?To inspect what the agent did and did not doObservability and provenance
Can I understand it?To make a better decision, not to read a longer answerVerification and defensibility
Can I influence it?To change a goal or stop an action without starting overSteerability and recovery
Can I own it?Authority matching the accountability they carryApprovals, escalation, decision rights, audit
Can I trust it appropriately?To know when to delegate, verify, supervise or refuseRisk-tiered autonomy and calibrated reliance

Left as vague aspirations, be transparent and keep a human in the loop, none of it can be shown to have been delivered. Written as the right-hand column, all of it can be tested.

Quiet Rejection

The most immediately useful part of the session is the list of behaviours, because most teams have all of them and count none of them.

People do not file complaints or refuse the tool. They keep the old spreadsheet, recheck every field, ask a trusted colleague to verify the answer, copy the output into another system and rebuild the reasoning by hand, and take screenshots for an audit trail the product does not preserve.

They use the agent for low-risk work and avoid it when the consequences matter.

From a dashboard all of that reads as adoption: a user signed in, a run completed, an output was accepted. From the user’s seat the workflow is now running the old process and the new one side by side, and the launch report looks healthy while the real work gets slower.

His human cost point deserves attention from anyone measuring this. Duplicate work creates fatigue, unclear ownership creates anxiety, and a polished answer with no inspectable evidence can make an expert feel their judgement has been reduced to clicking approve.

Comma

And the incentive trap that follows: reward visible usage while ignoring hidden verification work, and teams learn to protect the metric rather than improve the workflow.

One System, Two Lenses

He is careful that human experience and engineering response are not competing views. Human experience says what responsible participation requires; engineering makes those conditions durable rather than dependent on individual heroics.

Different roles need different proof. A legal reviewer may need citations and a record of which policy language applied. A security engineer needs logs, diffs, tool results and reproducibility. A business owner needs likely impact, a clear decision and a reliable escalation path.

The same agent can therefore feel transparent to one role and opaque to another, and that split appears inside a single team: the person deciding, the person operating and the person answering for the outcome each need a different view.

People also process differently. One wants a concise summary before choosing where to look deeper; another needs the sequence and the time to inspect it. One explores an unfamiliar interface comfortably; another is under pressure and needs the critical state to be unmistakable.

Note

Note: Test what your agents do, and what the people accountable for them can actually see. Try TestMu AI now!

The Human Contract

An agent never arrives in an empty workflow. A diagram shows steps, inputs, outputs and decisions, while the human system also holds expertise, exceptions, informal safeguards, relationships, status and responsibility, most of it undocumented.

Inserting an agent does not simply automate a box. It moves judgement, visibility and power.

An expert becomes an approver without anyone naming a role change. A coordinator becomes an exception handler. A manager stays accountable while losing access to the evidence they need to exercise judgement.

His remedy is a checklist per consequential step: name the actor, the authority boundary, the evidence required, the last reversible point, an escalation owner, and the state a human can actually observe.

His procurement example makes the granularity concrete. Suggesting three suppliers is one kind of action, drafting an order is another, and submitting the order, committing funds and notifying a supplier changes the authority and recovery requirements again. Calling all three agent completed task hides the part people are accountable for.

Readiness as Requirements

This is where the readiness figures come in, and he is careful about what they do not say: not that every implementation fails, and not that every employee needs to understand a model’s internals.

What they point at is uneven shared understanding of what has been implemented, how agents work and where responsibility sits.

When people do not know what a system can and cannot do, when it is acting, or whose authority it uses, trust becomes guesswork in both directions. Some overestimate the agent and accept too much; others underestimate it and duplicate work that no longer adds value.

His reframe is to treat readiness as a requirements problem rather than a communications footnote: put task-level boundaries into the experience, connect evaluation evidence to real work scenarios, and let people practise success, ambiguity and failure before the workflow becomes routine.

The test he proposes is sharp. Ask an accountable user to predict what the agent will do in an edge case, what it will not do, and where it will ask for help. If their mental model and the system’s behaviour diverge, that is a release risk. And a licence, a login or a completed training module proves access rather than readiness.

Can I See It?

The worked example runs through the rest of the talk, and he flags repeatedly that it is hypothetical: an incident response agent supporting an operations team during a service disruption.

It detects a disruption and recommends rolling back a recent configuration change, summarising it as a likely configuration regression with rollback advised. Concise, and it hides the work.

The incident commander needs to know which alerts fired, which service or region is affected, when symptoms began, what changed shortly before, which dashboards and runbooks were read, and whether the agent has already acted.

Gaps have to be visible too. Did a monitoring source fail? Is customer impact confirmed or inferred? If the agent used a five-minute-old deployment record beside a current error dashboard, the run view should show each source’s freshness, and it should reveal a failed dependency query, because that absence is part of the work.

His test for visibility is a person rather than a schema. Hand someone a completed run and ask them to reconstruct the path: can they identify the sources, tell what the agent changed against what it only proposed or failed to retrieve, find the first point where intervention was possible, and judge whether the evidence is fresh enough for this decision?

Explanation vs Verification

More explanation does not automatically create more understanding. A long answer can sound confident and be impossible to verify; a short one can be clear and omit the uncertainty that should change the decision.

His reframe is that the question is not whether the system can explain itself but whether the explanation helps someone make a better decision, and understanding is always attached to an action: accept, verify, edit, reject, escalate or investigate.

A good explanation separates four things: the conclusion, the supporting evidence, the uncertainty, and the next verification step.

He prefers reproducible artifacts over claims of invisible reasoning: show the diff, link the query, preserve the tool return, identify the runbook section and its version, show the policy result, and make independent re-runs possible. Hidden reasoning is not proof, and confidence in the wording is not confidence in the result.

When two sources conflict, the agent should surface the disagreement and offer a way to investigate rather than smoothing it into one confident answer.

His strongest testing instruction is to evaluate explanations when the agent is wrong, not when it is right. A polished explanation can increase reliance on a bad result when the citations are irrelevant, the diff is stale, or the verification step does not test the actual claim. An explanation that only supports acceptance is not supporting judgement.

Automate web and mobile tests with KaneAI by TestMu AI

Disagreement by Design

Disagreement is how expertise enters the system, not an exception to design around. If the only options are approve everything or start over, the workflow has not preserved meaningful agency.

In the hypothetical, the agent proposes rolling back across all regions. The incident commander agrees the change is suspicious but disagrees with the scope, wanting to test in one affected region, preserve traffic in a stable one, and hold customer communications until impact is confirmed.

Influence has to exist across the lifecycle: confirm intent before execution, distinguishing investigate a rollback from perform one; preview what will be read, created, changed, sent or deleted; require approval at a risk boundary for the exact proposed action; and let a person steer mid-run without discarding safe completed work.

Afterwards, support undo where undo is real and a compensating action where it is not, and do not use the same comforting words for both.

The backend has to make that control real, through idempotent tool calls, checkpoints, transaction identifiers, cancellation propagation, durable side-effect records and a known safe state. If an action crosses systems, cancellation has to mean more than closing a dialog.

His list of what to test is the ugly middle: a network failure after a side effect, a duplicate retry, stale authorization, partial success across two tools producing conflicting edits, cancelling while an external action is pending. A stop button that only hides the progress indicator is not control, and an edit box that changes the words but not the executable plan is not steering.

Ownership and Authority

If the agent acts and a human’s name is on the outcome, that human needs real authority, information and time. Otherwise human in the loop means responsibility without control, which he calls exposure rather than trust.

An interface saying only awaiting human approval hides the important question: which human, in which role, with what authority?

The roles to distinguish are recommender, approver, executor, monitor and escalation owner. One person may hold several, and the system still has to separate the functions.

Rules bind to task risk, data sensitivity, reversibility and blast radius, and the escalation is decided in advance rather than negotiated over chat while the service is failing. Querying monitoring may be automatic; changing one region may need the on-call commander; changing all regions may need a second approver.

Approval should authorise the exact proposed action for a limited time, and a material change should invalidate it. Change the target region, the rollback version or the affected service after approval, and the agent has to ask again.

Ownership also includes capacity and psychological safety. A named owner who is unavailable, overloaded or unable to understand the evidence is not a control, and in some cultures escalation reads as weakness, so whether the path is real depends on how leaders behave.

Calibrated Trust

He does not want to trust every agent equally and says you should not want him to. Trusting a system to summarise a low-risk document while refusing to let it send a consequential message is judgement, not inconsistency.

The goal is appropriate reliance: knowing when to delegate, verify, supervise or refuse, with trust varying by task, consequence, evidence and experience.

Autonomy is a risk control rather than a maturity badge, running from suggest, to collaborate, to act with approval, to bounded autonomy inside narrow permissions with monitoring and stop rules, each mapped to action risk and reversibility.

Raise it because evidence supports it across relevant conditions including ambiguity and partial failure. A first incident that ended well is not evidence.

Lowering it is part of the collaboration rather than a failure. If false alarms appear after a monitoring pattern changes, narrow the trigger and restore approval for that action without declaring the whole agent broken.

Both directions of miscalibration are worth watching. Over-reliance looks like rapid approval without opening the evidence; under-reliance looks like duplicate checking on work the agent does well, which may point at missing evidence, an unaddressed past failure, or a role carrying accountability without enough authority. His release gate is not that users trusted it, but that reliance matches demonstrated capability and task risk.

The First Visible Failure

The first visible failure teaches people more than the launch announcement did, because it shows what the organisation actually values when confidence meets reality.

Challenge the agent and get labelled obstructive, and everyone learns to stay quiet. Meet a bad output with did you use it correctly, and people learn that speaking up carries a cost.

Hiding failures to protect the project produces private workarounds that fragment the system: official metrics look clean while the real controls live in spreadsheets, messages and individual memory.

His version of psychological safety is not the soft one. It does not mean removing accountability or avoiding hard questions; it means people can surface uncertainty and disagreement without being punished for protecting the outcome, and people keep reporting when their concerns visibly become tests, boundaries or workflow changes.

The plan he leaves is deliberately small. Choose one consequential workflow rather than the whole organisation. Bring in the people who carry the consequences. Rather than asking whether they trust it, ask them to show how they work and how they verify, and treat every workaround they mention as evidence rather than as insufficiency.

Then run three scenarios, success, ambiguous and failure, with realistic data in a production-like environment, as a demonstration rather than a debate, ending by naming what is safe to release and what must limit scope. Log every no as a decision defect with severity, owner and exit criteria, and resist turning every gap into a tooltip: a stop action that does not propagate is an engineering defect, and an exception with no owner is an operating-model defect.

Q & A Session

Two audience questions closed the session.

  • What are humans rejecting that the tests did not catch, trust, tone or workflow fit?

    Justin Roy: All three route back to the five questions, so a rejection should be root-caused: was it not seen, not understood, not defensible? Workflow fit fails because agents get built from an isolated view of a job, and a requirements document rarely captures tribal knowledge or informal connections. His example is an asset switched off for 364 days a year that becomes critical at financial reporting, which no specification would have surfaced.

  • Have you ever had to kill a technically perfect agent because a human said no?

    Justin Roy: Yes, and he says the failure rate there is significant, because so much of what these systems do is a black box: a question and a data set go in and an answer comes out with no account of the logic behind the number. His case was a technically perfect agent serving an operational dashboard with every metric the team wanted, trending included, and no visibility into how the figures were derived, so the first person who arrived with a different number had no way to reconcile it. This is his own experience, offered without named organisations or figures.

This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.

Author

...

TestMu AI

Blogs: 232

  • Twitter
  • Linkedin

TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests