Next-Gen App & Browser Testing Cloud
Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Canal Mania and the Philosopher's Stone [Testμ 2026]
Canal Mania and the Philosopher's Stone [Testμ 2026]
Magnus Hillestad of Sanity on why agents are not digital employees, what Myriad cost at $10,000 a day, and the missing compiler for organising intelligence.

TestMu AI
Author
Published on:
Over one weekend an AI produced a full analysis of a company’s self-serve pipeline, good enough to hand to the finance team. The previous Wednesday, a team of agents asked to plan a single feature started planning, and kept planning, and never stopped.
Both happened at Sanity in the same week. In this session from Testμ Conf 2026, Magnus K. Hillestad, CEO and Co-founder of Sanity, uses the industrial revolution to explain why both are true at once, and what that says about where the AI revolution actually is. Nikhil Saxena, Brand Marketing Manager at TestMu AI, hosted.
If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.
TL;DR
Treating an agent as a digital employee slotted into an existing org chart is the wrong abstraction. Agents can be spun up for five minutes, cloned twenty times, programmed and checked programmatically, and the missing piece is a way to express intent to a team of them at a higher level than prompting.
- Why are agents not digital employees? - An agent does not need to fill a workday, can be parallelised or cloned, comes with something like a readme, and can have its context assembled, described and verified programmatically. None of that is true of the person whose seat it is supposed to take.
- Does more agents mean more output? - No. Sanity’s Myriad system gave some engineers more than 10x output, and adding agents produced more chatter and more meetings in their own workspace, with less finished work.
- What did running a fleet of agents cost? - Myriad could easily reach 4,000 dollars a day and on some days exceeded 10,000. A later design called Persona cut that by more than 90 percent for the same or better output.
- Did direct delegation solve it? - Not entirely. Magnus Hillestad shipped a production feature without writing code, and also went to bed with a perfect plan and woke to nothing done, because the agent delegated and nobody followed up.
- What do agents need that humans do not? - Clarity, context, checkpoints and correction, all delivered in a programmable way, because human delegation runs on informal interpretation that does not transfer to machines.
- What is the missing compiler? - The compiler is a metaphor for the absent abstraction layer between prompting one agent and directing a team of them, in the way assembly language sits below higher-level programming.
- Are today’s models the canals or the railroads? - Possibly the canals. Chips, data centres and models dominate the current frenzy, and Magnus Hillestad argues something can be foundational and transitional at the same time.
- Do you need to work at an AI lab to contribute? - No. The industrial revolution needed clockmakers, machinists and surveyors as much as the famous inventors, and the equivalent work in AI is organising intelligence rather than training frontier models.
He set out the contradiction before offering any resolution to it.
Capable and Failing
AI is astonishingly capable, and it fails spectacularly at many of the tasks you give it. His position is that both observations are true at once, and his own week supplied the evidence for each.
The weekend analysis worked. Sanity’s user base is growing extremely fast with revenue growing more slowly, and AI helped him produce an analysis of that pipeline solid enough to hand to his finance team to dig into further.
The Wednesday attempt did not. A team of agents asked to plan a new feature began planning and simply continued.
He put two public positions beside each other to show how far apart informed people are. As he characterised them, Dario Amodei holds that artificial general intelligence is one to three years away, while Alex Karp believes in AI and considers large language models completely and irresponsibly oversold.
Which leaves the question he spent the session on. Are we at the end of this, or the beginning?
The Philosopher’s Stone
To answer it he went back further than the industrial revolution, to people chasing something more ambitious than artificial intelligence. For centuries alchemists, Newton among them, searched for the philosopher’s stone.
It was never a product specification. It began as a stone, became a liquid, became other things, and what it really represented was ultimate capability, mastery over matter.
They never found it. What they developed instead was how to distil, crystallise, heat, cool, measure and combine substances, and once the myth fell away what remained was experimental knowledge and a new discipline: chemistry.
His parallel is that artificial general intelligence is similarly hard to pin down. His co-founder Evan expects we will get there; Magnus places himself closer to the Yann LeCun view, that this intelligence is not yet smarter than a cat and is nonetheless extremely capable.
He noted that CERN has technically turned lead into gold, and asked whether it mattered. His first lesson follows from that: we are defining the chemistry of intelligence on the way to AGI, whether or not we arrive.
Four Phases of a Revolution
For locating the present moment he used a heavily simplified version of Carlota Perez’s framework for technological revolutions, and was upfront that historians might not accept his compression of it.
- Eruption - which he dates from the transformer architecture through to the launch of ChatGPT.
- Frenzy - where he places us today.
- Crash - much discussed and not yet arrived.
- Deployment - the phase whose shape nobody yet knows.
England, 1790
The comparison he wanted is a specific year. In 1790 the steam engine could power a factory but could not move anything, the locomotive did not exist, there were no railroads, standardised parts had not arrived, steel could not be made cheaply, and the internal combustion engine was science fiction.
What Britain was doing instead was digging canals at extraordinary speed, the Bridgewater Canal being the first industrial one. Canals solved a real problem, moving coal, iron and stone faster and more cheaply than road could.
Private investors poured money in. Some of those investments were excellent, some foolish, some redundant, and some failed outright, which is why the period is remembered as a mania.
The canals worked, and almost nobody remembers them. What they left behind is the interesting part: engineers who had learned to survey, build locks and aqueducts, move enormous quantities of earth and coordinate projects at a scale Britain had never attempted.
The railways arrived later, in the deployment phase, and were built on that knowledge and on the navvies who had dug the canals in the first place.
Canals or Railroads
Applying that to now, the models, chips, data centres, networking and energy absorbing enormous capital may be our canals rather than our railroads.
He was careful not to dismiss them. Those things matter and we need them, and it is dangerous to assume that because they dominate the current moment they will define the final shape of the revolution.
On identifying the railroads in advance, his answer is refreshingly unbothered. You can try to work it out, or pretend to in order to seem clever, and it does not much matter either way.
The useful question is what we are learning from the canals. Something can be foundational and enormously valuable and transitional, all at the same time.
An Operating System for Work
His reading of the industrial revolution is that it was not a sequence of better machines. It produced a new operating system for work.
Before factories, most work happened in cottages and small workshops. Increasingly powerful machines created an organisational problem rather than a mechanical one: how do you arrange many people and machines around a shared production process?
The answers were innovations in their own right. Division of labour, standardisation of process, and systems for scheduling, supervision and quality.
Which makes the factory an organisational technology, and gives him the lens he applies to AI for the rest of the session.
The factory wasn't a building full of machines. It was an organizational technology a new operating system for work.
— TestMu AI (@testmuai) August 19, 2026
Magnus K. Hillestad on canal-mania and the philosopher's stone at #TestuConf26 pic.twitter.com/anvqGdWL1i
Not Digital Employees
The common move today is to insert AI into the existing human organisation. An AI software engineer, an AI researcher, an AI lawyer, an AI support agent, each taking a seat and standing in for the person who held it.
He thinks that is probably the wrong abstraction, because the differences run deeper than capability.
- No workday to fill - an agent can work for five minutes, then be spun down, which no employment model accommodates.
- Parallelisable - you can spin up twenty, and finally clone the coworker you always wanted to clone.
- Programmable - at least partly, where humans do not come with a readme file and use cognition to work out what they are supposed to be doing.
- Assembled context - deliberately constructed and programmatically described, and corrupted context can be discarded and rebuilt.
- Checkable by machine - the work can be verified programmatically rather than by someone eventually looking at it.
His related point about humans is the one people skip. Humans are not fungible and organisations differ by people, task, culture and context, which is why no single template runs a company well.
Note: If an agent’s work is going to be checked programmatically, the evidence has to be there to check. TestMu AI Test Intelligence tracks results, flakiness and failure patterns across runs, so verification rests on history rather than a single look. Try it free!
Myriad and Its Bill
The experiments are where the talk stops being theoretical. In December, after the release of Opus 4.5, his co-founder Simon called him unable to sleep, having built a system he named Myriad.
Myriad let a large number of agents work together in one place, and it looked rather like Slack. It handled memory better over time, and the more significant part was that it started from how humans and machines work together rather than from the model.
For some of their engineers and for Simon himself it produced more than 10x output on top of Claude Code and the tooling they already had. They released it publicly.
The cost was substantial. Runs could easily reach 4,000 dollars a day, and on some days went past 10,000.
The bigger problem was the net effect. More agents produced more chatter and more meetings inside their own little Slack, and less output, becoming talkative in a way he noted is rather human.
Persona and the Empty Morning
Simon’s next design, Persona, replaced the shared room with direct delegation from a human to an agent, letting agents collaborate on the back end to split the work and keep context small.
It worked, and the number is the headline: more than 90 percent cost reduction for the same or better output. Magnus shipped a feature into production without writing a single line of code, talking to an agent that worked with a team of agents.
He was equally direct about the mornings it failed. He would go to bed with a perfect plan and wake up to nothing moved, because the agent had delegated the work and nobody followed up, and eight hours of expected machine coding produced nothing at all.
They are now on the third version of the system, which is the honest shape of this work.
What he draws from it is that delegation becomes programmable, and that what we know about delegating to humans does not transfer. Human delegation is informal: a manager communicates intent, people interpret it, ask questions, report progress, and someone eventually checks the result.
Agents are worth thinking about as human-like, because they are trained on knowledge humans communicated. They also need four things humans do not: clarity, context, checkpoints and correction, each supplied in a programmable way.
Still Writing Assembly
His framing of where the tooling sits will land with anyone who writes code. We currently talk to agents at a very low level of abstraction: writing prompts, providing context by hand, splitting the tasks ourselves, deciding how checks happen, noticing failures and restarting work.
That, he suggests, may eventually look like programming in assembly language, the bridge between binary and the higher-level languages that came after it.
A mature system would let a person express intent at a higher level. It works today for simple linear tasks handled by one model, and falls apart when you try to abstract the work across a team of agents solving something complex over time.
He connects it back to management, noting that leadership is itself an abstraction upward, and that what is needed is a system for organising human cognition alongside machine intelligence.
Clockmakers and Surveyors
The closing argument is about who gets to build this, and it returns to the history one last time. We remember Watt, Stephenson and Arkwright, and industrialisation genuinely needed them.
It needed far more people than the ones who invented the headline machines. Clockmakers, blacksmiths, machinists, surveyors, civil engineers and navvies, most of them transforming existing knowledge into new practice by rethinking their own craft.
His term for them is people who understand one part of the system deeply enough to make the rest of it possible, and his claim is that the AI revolution will not be built only by those chasing the philosopher’s stone.
The open problems he listed are all organisational rather than architectural. How machines receive intent, how they divide work, how they coordinate without polluting each other’s context, how they know when they are finished, and how we check what they did.
Underneath sits the one he finds most interesting: how leadership gets reinvented when humans and machines work together, and what language describes the organisational structures that result. Obvious in twenty years, unknown today.
His four lessons close the loop. We are defining the chemistry of intelligence rather than chasing the stone; you do not need to predict the next railroad to have impact; we need a system to organise human cognition with AI; and building around the core technologies is not cleanup work, it is the revolution.
Q & A Session
The Q&A is worth reading for what he declined to answer as much as what he did.
- When an agent produces different outputs for the same input, how do you distinguish acceptable non-determinism from an actual defect?
Magnus K. Hillestad: He said plainly that he does not have the answer, and told the questioner to go and work it out. What he does believe is that the intersection of the probabilistic and the deterministic matters enormously, and Sanity is working on structuring content so that probability mixes with deterministic knowledge, because a wholly probabilistic approach is not enough. He flagged his own bias in saying so.
- How do you design prompts that let an agent ask clarifying questions without frustrating the user?
Magnus K. Hillestad: No silver bullet, and the answer depends on the craft and much more on the abstraction level. His example is internal: Simon, their CTO and deep on architecture, has the time and inclination to sit splitting tasks with the machines, where Magnus as CEO has less of both and needs more broken down for him. He turned it back into a human question, asking which follow-ups frustrate colleagues in any organisation, and concluded you have to define the context of how work is done rather than only the context of the answer. He cited Trillion Dollar Coach on Bill Campbell: faced with a problem, the first question is not what the answer is but who should solve it.
- How do you test a voice agent’s prompt, at audio or transcript level, and how should accessibility factor into voice prompt design?
Magnus K. Hillestad: He turned both down. Voice agents are not something his team has spent time on, and rather than improvise he said he had nothing useful to add and redirected to questions about systems and the delegation of work.
- Which elements of machine intelligence are least compatible with human intellect?
Magnus K. Hillestad: His light answer was time, which he admits annoys him. Machines cannot understand time, and because agents are trained on human material they think in time while operating on an entirely different footing, which makes planning hard. His more considered answer is the context window. Agents should be much narrower than we tend to build them, and there is a real difference between training a capable model and operating the context around one, which is where most of his team’s work goes.
This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.
Author
TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests



