Next-Gen App & Browser Testing Cloud
Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Scaling Enterprise Practice in the Agentic Era [Testμ 2026]
Scaling Enterprise Practice in the Agentic Era [Testμ 2026]
Five enterprise quality leaders on why pilots stall: capacity is no longer headcount, autonomy must be earned through evidence, and testing becomes trust work.
Published on:
Most agentic pilots do not fail because the agent underperforms. They fail because the organisation around it cannot absorb what the agent produces.
That is the thread running through this Testμ Conf 2026 panel, and every one of the five panellists arrives at it from a different direction.
Speaking are Mhahesh Muraleedhara, Head of Quality Intelligence for North America at Zensar Technologies; Dror Avrilingi, Amdocs Studios CTO and Head of Quality Engineering Studio at Amdocs; Vikul Gupta, Chief Technology Officer at QualityAI; Richa Agrawal, AVP and Senior Principal Architect for Quality at GlobalLogic; and Nagendra BS, Senior Vice President for DevOps, SRE and Digital Assurance at Hexaware.
If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.
TL;DR
The constraint on scaling agentic engineering is organisational rather than technical. The unit of capacity stops being headcount and becomes human expertise plus agents plus reusable intelligence, which breaks the hourly commercial model, the testing pyramid, the definition of done and the apprentice career path. What turns a pilot into a practice is reliable context, verification designed in from the start, a named owner after deployment, delivery budget rather than innovation budget, and autonomy earned through evidence.
- Does scaling agentic engineering mean replacing testers? - No. Richa Agrawal calls that a boardroom myth and says agents assist with human involvement. Dror Avrilingi sharpens it: the change is not replacing people with agents, it is replacing headcount as the unit of capacity.
- What assumption broke first? - That capacity equals headcount. Dror Avrilingi argues the unit becomes human expertise, the agents, automation and reusable intelligence, and that once it changes the testing pyramid shifts upward and services pricing has to become outcome-based.
- Will AI change how services firms price work? - Yes. Nagendra BS expects the hourly rate to split into one component for humans and one for AI tokens, bundled into a single price, and is explicit that no vendor can raise an invoice because tokens cost more than expected.
- What did agents produce in the department store programme? - Mhahesh Muraleedhara reports a quarter of planning carrying 1,500 epics and 15,000 user stories, where test scenarios reached full traceability, roughly 80% of test cases were directly consumable and automation code was 70 to 80% usable, finished before the planning exercise itself was.
- What broke once the agents got fast? - Test data provisioning. Mhahesh Muraleedhara says agents generated faster than any provisioning process could feed them, data became the queue and environments sat immediately behind it, so the definition of done had to change.
- Is functional coverage enough for an agentic system? - No. Nagendra BS found on an airline chatbot programme that the completeness of the agent workflow mattered more than the correctness of individual answers, which is what pushed the team from functional testing towards AI assurance.
- What are the five signals a pilot will reach production? - Vikul Gupta names a use case tied to a real workflow and a measurable outcome, reliable context, verification designed in from the beginning, a clear operating model with a named owner after deployment, and a plan for orchestration, observability and governance beyond the pilot.
- Is an impressive demo a strong signal? - No. Dror Avrilingi says he is less interested in how impressive the agent is than in how the organisation behaves around it, and rates willingness to delegate the activity in production as the clearest predictor.
- How does the budget reveal whether a pilot is serious? - Innovation budget means it is still a demo; capital delivery budget means someone has decided to get the work done. Mhahesh Muraleedhara says the source of the money tells you more than the amount.
- Should every team get the same autonomy? - No. Vikul Gupta prescribes a hub-and-spoke model with a shared enterprise foundation, where teams start in assisted mode and move to on-the-loop for well-bounded tasks, because autonomy should be earned through evidence rather than granted after one successful pilot.
- Should a platform be built per use case? - No. Richa Agrawal declines to build agents per use case and builds for the complete workflow, citing a customer with roughly 5,000 regression tests who wanted a single script-generation agent, and notes that proofs of concept typically cover only the happy path.
- Where does the panel think QE lands? - All five expect the discipline to stop being called testing. Dror Avrilingi predicts chief quality officers and trust engineers owning an enterprise trust index, and Vikul Gupta names the successor practice continuous trust engineering.
Capacity Is Not Headcount
Richa Agrawal opens by naming the myth she meets in boardrooms, which is that agents replace people. Her position is that agents assist with human involvement, and that the confidence for anything more is not there yet.
Her diagnosis of why scaling stalls is that teams keep building use cases and leave the operating model behind.
Dror Avrilingi identifies the specific assumption that held for years and no longer does: capacity equals headcount. In the agentic era the unit becomes human expertise, the agents, automation, and reusable intelligence brought into an engagement.
Two consequences follow immediately in his account. The testing pyramid shifts, with everybody moving up it, and the services commercial model has to move toward outcomes.
His formulation is the cleanest correction on the panel: this is not about replacing people with agents, it is about replacing headcount as the unit of capacity.
The Hourly Model Breaks
Nagendra BS grounds that in how services firms actually bill. Everything is priced on hours, whether it is packaged as fixed price or fixed capacity, at some rate per hour, and optimisation historically meant less effort and therefore a smaller bill.
His forecast is that the rate splits in two: one component for humans and one for AI tokens. What matters is what he says next, which is that the customer should not have to care about the split, so suppliers have to bundle both into one price.
He is explicit about the thing vendors will be tempted to do and cannot. You cannot go back to a client and say tokens consumed more cost so the invoice has gone up.
Pricing models and operating models therefore have to absorb it, and his view is that the industry should accept that and move on.
1,500 Epics, 15,000 Stories
Mhahesh Muraleedhara brings the panel’s one concrete programme, at a large American department store, where a single quarter of planning carried 1,500 epics and 15,000 user stories.
The agentic quality engineering workflow produced test scenarios with full traceability, test cases he puts at around 80% directly consumable, and automation code at 70 to 80% usable, all completed before the quarterly planning exercise had finished.
These are vendor-side numbers for an unnamed client, self-reported rather than independently measured, which is worth holding in mind alongside his own comment that saying it six months earlier would have got him laughed at.
What preceded the result matters as much as the result. The team spent time establishing the foundation for agentic adoption inside quality engineering first, with the stated goal of making QE a front-runner rather than the function everyone perceives as the bottleneck.
The Next Bottleneck
The most useful part of his account is what went wrong, which arrived within weeks.
Test data provisioning became the binding constraint. Agents generated faster than any provisioning process could feed them, so data became the queue, and environments were immediately behind it. His phrase for the second one is the same story with a different resource.
The definition of done had to change, and clearing each bottleneck exposed the next.
His conclusion is the transferable one. The team was not accelerating the pipeline so much as discovering what the pipeline actually was, and most of it had never been written for an agentic cycle. Producing assets is pointless if the sprint cannot consume them at the rate they arrive.
Note: Give your agentic pipeline an execution layer that scales with what it produces. Try TestMu AI now!
AI for IT, AI for Business
Nagendra BS splits adoption into two buckets that are worth separating in any conversation about ROI: AI for IT, where a provider delivers development, testing and DevOps more efficiently, and AI for business, where the end customer makes their own processes more efficient.
His case study is both at once. An airline customer built a customer support chatbot covering queries across 12 areas, including baggage claim, real-time flight cancellations, booking changes and customer profile.
The architecture was genuinely complex: retrieval-augmented generation alongside FAQ retrieval, multi-turn conversation handling, and integrations with almost 15 backend systems including a rules engine and a channel proxy.
Testing started conventionally, with functional coverage of happy paths and edge cases area by area. The team quickly found that was not enough, because the biggest issue was the completeness of the agent workflow rather than the correctness of any individual answer.
Testing What Agents Should Not Do
His reframing of the testing problem is the sharpest line in that stretch of the panel: it is not only testing what agents are supposed to do, it is testing what they are not supposed to do.
He reaches for two reference points, both recounted from memory rather than cited. One is a reported incident in which an agent under evaluation found a route out of an environment it was not meant to leave, which he treats as the industry’s watershed example of an agent going wrong in an unintended direction. The other is a benchmark that built a simulated enterprise and tested processes across departments, where he recalls the best available models reaching only around 35 to 36% accuracy end to end.
Neither is named precisely enough in the session to reproduce as a citation, so both are best read as the shape of his argument rather than as evidence you could quote onward.
The practical obstacle he did have to solve in delivery was the observability layer, meaning which agent-specific logs can actually be analysed.
His reasoning for why that is hard holds regardless of the examples. Traditional code is deterministic and a human wrote it, so the outcome is expected. In an agentic system decisions happen in real time and you do not know which agent is calling which, or what influence they have on each other, so you need a real-time view rather than a post-hoc one.
His prescription is continuous autonomous validation of agents, running alongside conventional functional testing rather than replacing it, covering decisions, policy adherence, tool usage, prohibited actions and final state. He expects agents to end up monitoring agents, on the grounds that agents work around the clock and people cannot.
"We aren't just supposed to look at what agents can do, but at what agents are not supposed to do."
— TestMu AI (@testmuai) August 20, 2026
A sharp reminder from Nagendra BS, Senior Vice President, DevOps: the biggest early miss with AI was overlooking the agent workflow itself. The fix was to start testing the AI, not… pic.twitter.com/Psgs29IGMU
Five Signals for Production
Vikul Gupta answers the pilot-to-production question with five things to look for. He gestures at analyst reports his clients quote at him, without naming any, and offers the signals as the more useful test.
- The right use case - tied to a real workflow or a measurable business outcome, with a clear owner and economic value. A programme that only demonstrates an agent generating something impressive stays a pilot.
- Reliable context - the one he says fails most often. Requirements, existing test cases, coverage, trusted data, business rules, permissions, historical knowledge and the integrated ecosystem all count, and weak context is among the fastest ways to lose a pilot.
- Verification designed in from the beginning - knowing how outputs and actions will be evaluated, what evidence is retained, what the thresholds are, and when a human is escalated to.
- A clear operating model - somebody owns the agent after deployment, including performance, cost, security and exceptions. Where ownership stops at the proof of concept, the programme usually stops too.
- A plan beyond the pilot - orchestration, observability, governance, model access and continuous context improvement.
His summary test compresses all five: does the pilot prove that the technology works, or that the enterprise can operate it safely, economically and repeatably?
Dror Avrilingi answers the same question by moving the subject. He is less interested in how impressive the agent is than in how the organisation behaves around it, and wants an outcome that is measurable rather than semantic, work embedded in the existing workflow rather than in a separate tool, and above all a willingness to delegate the activity to agents in production.
On embedding he is blunt about the failure mode: if someone on the ground has to step out of their workflow into another AI tool, the value is lost, and that is where pilots die.
Sponsor, Budget, Metrics
Mhahesh Muraleedhara offers three tests, and all three are about the organisation rather than the technology.
- The sponsor - does the person leading the implementation have authority to change the definition of done? Without it the pilot never becomes a process; it becomes an excellent readout, a pat on the back, and then nothing.
- The money - innovation budget means the organisation is still thinking about production, while capital delivery budget means somebody has decided to get the work done. The source tells you more than the amount.
- The metrics - is anyone measuring things that could embarrass them? He has sat in ceremonies where the only concern was whether sprint velocity took a hit. Programmes tracking defects, rework, reviewer load and defect escape are the ones managing a real transition, because those are the numbers that can go wrong fast.
Hub and Spoke Autonomy
Vikul Gupta borrows multi-speed IT from the automation era to make a point about uneven maturity: teams were not all at the same level then and are not now, across process maturity, people maturity and tool exposure.
Step one is a shared enterprise foundation, the hub, covering agent access, orchestration, security, observability, evaluation, knowledge management, context creation and auditability.
His reason for centralising is concrete rather than architectural preference. Let federated teams build those foundations independently and you get fragmented architectures, inconsistent controls, and nothing auditable, because no standard exists to audit against.
Step two is graduated autonomy. Lower-maturity teams start with agents in assisted mode, helping with test design and cases or helping manual testers prioritise, and as a team demonstrates performance and stronger controls the agents move to on-the-loop for well-bounded tasks.
The principle he states for it is the one to carry into any rollout plan: autonomy should be earned through evidence, not granted because a pilot succeeded. His warning about the alternative, that almost everyone who tries to flip the switch outright fails, is rhetorical emphasis rather than a measured figure, but the direction is clear enough.
Governance is federated to match: the hub defines guardrails, monitoring standards, auditability and minimum evidence, while project teams stay accountable for the use case, the domain context and the outcomes.
Platforms, Not Use Cases
Richa Agrawal reports a pattern worth noticing: customers who never got conventional automation working now want an autonomous intelligent layer across the whole estate.
Her answer is a commercial position rather than a technical one. Her team declines to build agents for individual use cases and builds the platform for the complete workflow instead.
Her example of the ask she turns down is specific: a customer with a regression suite of roughly 5,000 test cases who wants one agent to generate test cases faster.
Her critique of typical proofs of concept is that they cover only the happy path, without modelling what the failure is, where it occurs, or how the system learns from it and recovers.
What the platform ships instead is a set of trust signals established before anything scales: outcome accuracy, consistency, success rate, error rate, false positives and negatives, and predictability. The second layer is traceability of agent decisions, why a decision was made, whether it was right and which agent made it, driven by test results, defects, risk and business rules. Rollout starts with one product or one tower and widens once those hold.
From Testing to Trust
The moderator closed by asking where enterprise quality engineering lands in three years. All five answers converge, which is unusual for a panel.
Dror Avrilingi refuses the three-year horizon and proposes six months. He expects organisations to have a chief quality officer whose engineers split between AI for QE, using AI to verify faster, and trust engineers doing QE for AI.
The one metric he says that officer must own is an enterprise trust index that tells the CEO whether the agents are working, the workflows are talking to each other, and whether a release diluted the AI backwards.
Mhahesh Muraleedhara raises the human cost, and it is the point the others do not make. The traditional pyramid put the heavy load on the junior layer, which was also the apprentice model where people learned how systems actually work. Agents taking that load removes the training path, and he expects a diamond rather than a pyramid: fewer, more senior people with a platform that gets them to outputs faster.
Nagendra BS deliberately takes a different angle, saying he has no doubt the technology works and that the open question is the economics, which he expects to follow the same path as cloud. If the last decade was automation for speed, he calls the next one engineering trust.
Richa Agrawal expects the industry to stop using the word testing. The QA engineer becomes the person who defines what good looks like, what risk is acceptable and what agents are allowed to do, and the question changes from whether testing is complete to whether there is enough evidence to trust the release.
Vikul Gupta, going last, would remove testing and QA from the vocabulary entirely and call the practice continuous trust engineering, with repeatable work going autonomous and teams left on system intent, risk modelling, evaluation and governance.
On assurance he adds two practical things. Explainability, in his view, is largely a matter of using the right libraries to extract what a model already exposes. And for judging output, run the same work through a second or third model and bring a human in when the variance grows, with metrics grouped into a handful of scores rather than dozens, so a customer is not asked to interpret forty numbers.
This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.
Author
TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




