Next-Gen App & Browser Testing Cloud
Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Will the Real Autonomous Agent Please Stand Up [Testμ 2026]
Will the Real Autonomous Agent Please Stand Up [Testμ 2026]
Dona Sarkar on why AI is stuck in its scroll bar moment, the three kinds of software nobody knows how to test, and the faster horse test.
Published on:
Software whose customer is an agent rather than a person. Apps that fork a billion ways because every user carries different memory. Runs that cannot be reproduced because the agents that did the work have finished and vanished.
In this keynote session from Testμ Conf 2026, Dona Sarkar, Chief Troublemaker for Microsoft Enterprise AI Advocacy at Microsoft, named those as the three things nobody has worked out how to test, and she includes herself.
If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.
TL;DR
Agent-facing software is software whose customer is an AI agent rather than a person, reached through API calls where the interface stops mattering. Dona Sarkar names it as one of three kinds of software nobody has worked out how to test, alongside apps that fork per user and runs that cannot be reproduced.
- What is the faster horse test for an AI project? - Dona Sarkar’s rule is that if your AI project would make sense as a 2016 startup pitch, it is a faster horse. She adds explicitly that this is okay, while pointing out you are then competing with yourself yesterday rather than creating a new way to do anything.
- Why is AI harder to reimagine than mobile or VR? - Because there is no new form factor. Dona Sarkar points out that the internet brought a browser, mobile brought a phone and her hologram work brought a headset, each a physical thing that invites the question of what new problem it could solve. AI arrives without one, so teams default to porting.
- Which kinds of software does Dona Sarkar say nobody knows how to test? - Software with no interface, where an agent is the customer and the design no longer matters because it works through API calls. Apps that fork a billion ways because each user’s memory and context differ. And mortal software, where an agent spawns agents that finish and disappear, so the run cannot be reproduced.
- Does she offer a way to test agent-facing software? - No, and she says so directly rather than implying it. Asked whether we know how to test for a world where the user experience does not matter because an agent is the customer, her answer is that she does not, and she repeats the same admission about the autopilot she uses daily.
- Will AI take testing jobs, in her view? - Not in the way the headlines put it. Dona Sarkar notes a robot-apocalypse headline ran back in 2019, seven years of the same conversation, and says robots are not exactly here to take our jobs. Her sharper claim runs the other way: we need AI to take some of your jobs, meaning the repeatable parts.
- What is the difference between a copilot and an autopilot? - Copilots are reactive and act when you call them. Autopilots run in the background with their own identity, take action on your behalf and tell you when something needs doing. Dona Sarkar says Microsoft announced autopilots at its May developer conference, with the first called Scout.
- Was Scout demonstrated? - No. Dona Sarkar describes Scout giving her hourly updates in Teams and a daily pass over external mail, telling her which meetings lack artifacts and asking whether to move one. Nothing was shown on screen, and she is a Microsoft employee describing her own employer’s product.
- How should testers split their work with AI? - Using her fashion analogy. Off-the-rack work, meaning first drafts, formatting, cleanup, data entry and first-round research, goes to AI. Tailored work you do alongside AI while staying in the driver’s seat. Couture work, meaning writing, storytelling and inventing new ways to test, you hoard.
- What is the PI framework? - People, ingredients and experiments. Catalogue hidden talents and side hustles rather than the org chart, inventory skills, relationships, internal technology, customer access and unused data, then ask what problem you could not solve last year that you might solve now and run experiments against it.
- How many AI experiments should a team run? - Dona Sarkar gives three different numbers in one keynote. Her own team runs a hundred, she then says you do not have to and she would do about ten, and her homework for managers is twenty-five across the autumn. Treat it as a range rather than a prescription.
- What new job titles has she seen appear? - Testing agent boss, which is what she calls herself, reviewing fleets of testing agents. Plus guardrail guru for limits, cost and permissions, context captain for the right data and steering, taste titan for design judgment, experience director for flow, and intent translator. She says these emerged at Microsoft and some customers rather than from research.
- How do you reduce hallucination cascades in autonomous agents? - Asked specifically about reflex loops and multi-tiered memory, Dona Sarkar addresses neither. She offers three operational defences instead: very clear instructions as the first line, old-school authentication and authorisation so an agent does not inherit your permissions, and logging everything early. She adds that an agent will not work until roughly five runs.
Seven Years Of Apocalypse
She opens on a slide showing a newspaper headline about AI writing just like the author and bracing for the robot apocalypse, and runs a guess-the-year poll. The audience works backwards through 2026, 2025, 2024, 2023, 2022. The answer is 2019.
Her arithmetic from that is the point: the industry has been talking about the jobs apocalypse for seven years, which she calls an insanely long time to fixate on one topic. The article’s author, date and link are never given beyond the year.
Her verdict on those seven years is hedged rather than dismissive. Robots are not exactly here to take our jobs, she says, and the qualifier matters.
She then describes the corporate phase that followed, where every chief executive decided the company needed AI for everything, bought every product, and put leaderboards up so the heaviest users got bonuses.
That turned out badly in her telling, because people used it for genuinely silly things and it became expensive. She dates the wave without a source: token maxing started earlier in the year and died out over the summer when the bills arrived. No company, bill or spend figure sits behind it, so it is her characterisation of an industry mood.
The Faster Horse Test
She grew up in Detroit with both parents working at Ford, and quotes the line from the company wall about people asking for faster horses if you had asked them what they wanted.
She then debunks it herself in the same breath, saying Henry Ford did not actually say it and it is just a cool quote. Anyone reproducing the quote without her correction misrepresents her. She also asserts in the same passage that Ford invented the car, which he did not.
Her gloss on the idea is that most people do not know what they want, and try to innovate by doing the thing they already do, faster.
The slide pairing she distils it into: faster horses optimise the present you are in, and cars create a better future that does not exist.
The concession inside that line is worth keeping. She is naming faster horses, not condemning them, and 2016 is her rough pre-LLM marker rather than a defined threshold. She applies the same test to her own record, saying the industry ran plenty of faster-horse exercises across her 22 years.
Mailing Discs
Her operating-system anecdote is told from memory and undated. The team wrote C++ that compiled for a day and a half in a lab, then someone took the bits, burned them onto a disc, put it in a box and mailed it to your house. That is how software got tested.
The consumer half of the loop was you inserting the disc and testing it, which she calls an insane thought now, when a younger colleague would ask what a CD-ROM even is.
The faster-horse fix would have been building more build labs, buying planes and mailing discs to people faster.
The car version asks why anything is being mailed at all, and lands on issuing a download to people’s machines in minutes rather than running a month-long shipping project.
The numbers in that story, the compile time, the number of labs, the minutes and the month, are all illustrative and carry no source or baseline. The point for testers is that what got reinvented was the delivery and testing model, not the speed of any single step.
The Scroll Bar Moment
She shows her own graduation photograph holding a flip phone and lists the three things it did: calls after a certain hour, texting slowly on a number pad, and buying ringtones.
The core observation follows. When the first iPhone arrived it tried to behave like a phone, and the first apps had scroll bars, because they were web browsers made small with a scroll bar put on. Was that great? It was not, and nobody knew better at the time.
The generalisation she wants readers to keep is that every era has a scroll bar moment, because people always port before they reimagine, and that is where the industry sits now.
Her patient counter-example is Netflix mailing discs, gathering a customer base, and being positioned to reuse it once internet speeds caught up. She dates the founding to 1998; the company was founded in 1997.
The implied test for a quality reader is whether the AI work in front of them is a shrunken web page with a scroll bar bolted on, or an interaction that could not have existed before.
Note: Ask whether you are porting the old thing or inventing a new one. Try TestMu AI now!
No New Form Factor
Her diagnosis for why AI resists reimagining is that it has no new form factor. The internet brought a browser, mobile brought a phone, and her hologram work brought a thing that sits on your head. AI brings none of those.
Her counter-example is Shazam, credited to an engineer she names and describes as a Microsoft intern with a good music collection whose colleagues kept asking what was playing. The surname in the captions does not match the publicly known co-founder and the intern detail is uncorroborated, so no name is printed here.
The argument that matters survives the sourcing problem. Shazam did not exist on a laptop before, and it was not ported. It was a new thing that had never existed.
Her technical justification is wrong and she half-corrects it mid-sentence, starting on accelerometers before switching to the ability to hear music. Song recognition is a microphone problem, and the corrected version is the one that stands.
She adds two follow-on reimaginings: run tracking, where the alternative used to be leaving string behind you, and ride hailing, where you drop a location and say come and get me here.
Replacement Is The Wrong Frame
The framing she rejects is the one everyone repeats: that AI will do your job better than you, that it will replace testers, that it will replace developers.
Her rebuttal is the cleanest line in the keynote. Nobody ever said cell phones were going to replace laptop users, and saying it now sounds absurd, because the two were not trying to do the same thing.
That is the pivot from history to prescription. Everything after it concerns which assumptions and which artifacts change once you stop porting.
The Assumptions AI Unlocks
She sets up by listing the mobile assumptions that unlocked mobile thinking: the thing is vertical, pocket-sized, always connected, touchscreen first, has a camera, knows where you are, and can interrupt you.
The AI equivalents she reads off a slide: always connected, any form factor, knows the world’s information at all times, runs in the background, needs no greeting, speaks your language, and can take actions on your behalf.
Her payoff is that once you accept those, the set of problems you can solve and the things you can invent become different from what they were.
The testing surface she says will have to be covered goes well past browsers and smartphones, taking in head-worn devices, a command line, watches, ambient speakers and local automation. One product name in that list is not recoverable from the captions and is left out.
The one she explicitly parks is robots, where she says the era has not even started. That list is a set of surfaces rather than a test strategy, and she offers no tooling, harness or coverage model for any of them.
Three Untestable Kinds Of Software
The first is software with no interface, built for agents. Her worked example is a code host written originally for human developers, with a UI to match. Agents do not work that way and do not care whether the design is good, because they work through API calls.
The question she leaves open is what it means when your software’s design no longer matters because an agent is the customer, and she answers it plainly: do we know how to test for that, because she does not.
The second is the app that forks a billion ways. The history, memory and context an AI product holds on one user against another are completely different, so two people can use the same product and have entirely different experiences.
The third is mortal software. You start an agent, it spawns other agents, they do their work and then they are gone, so the run cannot be reproduced. Her illustration is code that rewrites itself around where its user happens to be.
The upside cases she reaches for are hedged and unsourced, including a heavily qualified aside about pattern-finding in medical research that she stumbles into and does not stand behind. It is not reported here as a result.
Her closing plea in this section is for patience, noting how long it took for a phone to become a boarding pass, a wallet and a primary source of information, and that nobody yet knows how to build and test things that are natively AI.
Chat, Artifact, Auto
Her analogy for where this ends up is spell check. You used to go to a menu and run it, then squiggly underlines appeared, and now it happens without anyone thinking about it. That, she argues, is the direction: invisible.
The three phases she names are chat, where you ask an assistant something and it answers; artifact, where you ask it to make something with you; and auto, where it runs in the background and acts on triggers.
A revealing aside about the deck on screen: she made the slides herself and had assistants try to improve them, which she offers as the reason they look random. The second product name in that sentence is garbled and is not reproduced.
The direction of travel she wants remembered runs from her prompting AI, to her prompting AI to prompt you.
Copilots are reactive and act when called. Autopilots run in the background, have their own identity and take action on your behalf. She says Microsoft announced autopilots at its May developer conference and that the first is called Scout.
Her Scout usage is described and never shown: hourly updates internally, a daily pass over external mail, telling her about the keynote and a customer meeting and which one she lacks artifacts for, and asking whether to move something. Then, three times over, that she does not know how to test it.
First you ask AI for answers. Then you ask it to build. Now it runs in the background and acts before you ask. Catch Dona at TestMu Conf 2026 on the chat-artifact-auto shift, and why the AI-native apps worth building need your expertise, not just a bigger model. pic.twitter.com/KHN7Ccv0LC
— TestMu AI (@testmuai) August 21, 2026
Off The Rack, Tailored, Couture
Off the rack is what you buy in a store, tailored is what you have taken in to fit, and couture is made from scratch. She maps the three onto kinds of work.
Off the rack goes to AI: the things you have done a thousand times, keeping up with the industry, watching what is happening in your repositories, first drafts, formatting, cleanup, data entry and first-round research. She expects it to do those better than she would.
Tailored is collaboration with you still driving, working through which updates apply to your work and whether you have the access, licence or subscription to prototype against them.
Couture is what you hoard: writing and delivering content, inventing new ways to test software, and finding better ways to tell stories.
The forcing function is that you cannot add work without shedding some, and doing work AI can do means competing with yourself. That is where her actual answer to the job question lands: we need AI to take some of your jobs. The instruction is concrete, to list all your work, sort every item into one of the three buckets, hand off the first, tailor the second and hoard the third.
The PI Framework
Her counter-evidence on job loss is a single anecdote with no source, that a major AI lab hired a head of recruiting the previous month, and her question of whether all the jobs are going away when everyone is hiring recruiters.
The mindset she attacks is the finite pie, where leaders want a bigger slice, which she calls an amateur way of thinking. The pie on the slide is one she baked herself and cheerfully calls hideous.
People comes first: gather your people rather than the org chart, and ask about hidden talents, unusual hobbies and side hustles, plus what they know how to do that nobody there knows about.
Ingredients is an inventory, because she says people undervalue what they already have: skills, relationships, internal technology, which customers you can reach, and what data is sitting unused. The question those feed is what problem you could not solve last year that you might solve now.
Experiments comes with three different counts in one talk. Her team runs a hundred, she then says you do not have to and she would do about ten, and the homework sets twenty-five across the autumn. A thirty-day framework is mentioned once and never explained, so none of these is the recommendation.
The jobs she says have already appeared at Microsoft and some customers are testing agent boss, which she claims for herself as someone who reviews fleets of testing agents; guardrail guru for limits, cost, permissions and security; context captain for the right data and steering; taste titan for design judgment; experience director for discovery, navigation and exit; and intent translator, on the grounds that feeding AI human nonsense produces nonsense.
Q & A Session
The host said there were many questions and time for one, which she selected and read aloud. No audience member spoke on air and no asker was named.
- Which architectural mechanisms, such as reflex loops and multi-tiered memory consolidation, best minimise hallucination cascades in truly autonomous agents?
Dona Sarkar: Very clear instructions come first, and that is your first line of defence. Then the ironically old-school part, authentication and authorisation: an agent should not simply hold your permissions, and read access and write access are two very different things. Then log everything, especially at the beginning, though I do not mean forever: what access it had, where it got stuck, what answers it gave and where they came from. Hand an agent a pile of conference submissions to shortlist, then audit the process it used. And do not expect an agent to work until it has run about five times. Treat it like a video game, where you cannot move on until you clear a level.
Her final instruction to the audience was to post about the session themselves, and to write it rather than have AI do it.
This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.
Author
TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




