Hero Background

Next-Gen App & Browser Testing Cloud

Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

Next-Gen App & Browser Testing Cloud
Testμ

Reinventing the QE Practice at Global Scale in Agentic Era [Testμ 2026]

Four delivery leaders from Wipro, Capgemini, Nagarro and Persistent on guardian agents, the new QE pyramid, and where humans still sit in the loop.

Published on:

For two decades, quality engineering at global scale ran on one equation: more scope meant more testers. That is the premise the host puts to the panel in the opening minute, and it is the panel’s job to say what replaced it.

Nothing has replaced it cleanly. Over 64 minutes, four delivery leaders describe a pyramid re-forming around different roles, a set of narrow tasks agents now do unsupervised, and one problem none of them claims to have solved: proving the machine was right.

At Testμ Conf 2026, Raghavendra Prasad MG, Head of Sub Practice, AI/GenAI Solution Building for Quality Engineering at Wipro; Khimanand Upreti, Managing Director and Head of AI in Run BU at Nagarro; Jeba Abraham, Group Vice President at Capgemini; and Ninad P Shirodkar, VP, Solutions Consulting - BFSI & Quality Engineering at Persistent Systems, worked through what replaces it.

Youtube thumbnail

If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.

TL;DR

A guardian agent is a second agent run in parallel with a working agent, scoring its output continuously against metrics such as hallucination, toxicity and relevancy. It exists because an enterprise cannot prove an autonomous agent behaved correctly without something watching it that is not the agent itself.

  • Does agentic AI mean quality engineering needs fewer testers? - Not exactly fewer, but fewer added. Khimanand Upreti of Nagarro says headcount may still rise while the rate of increase falls, and keeps his own hedge that some situations will still need more people. Jeba Abraham of Capgemini says the dev-to-QA ratio has changed significantly and pointedly refuses to attach any number to it.
  • What does the new QE team structure look like? - Ninad P Shirodkar of Persistent Systems replaces the old tester, test lead and test manager pyramid with three layers: engineers and developers who use AI daily, SDETs and AI engineers who build the platforms and agents, and AI architects who compose and orchestrate agents and handle context setting. He treats the middle layer as a genuinely different skill set rather than renamed automation work.
  • Which testing tasks do agents already do without a human? - Ninad P Shirodkar names requirement impact analysis, test case and scenario creation, and UI and API script generation. Raghavendra Prasad MG of Wipro offers four more as examples rather than a closed list: requirement quality, scenario identification, synthetic data generation and self-healing agents.
  • Which testing work still needs a human in the loop? - Ninad P Shirodkar keeps humans on test data conditions, all reviews and test data management. Raghavendra Prasad MG draws the line at business criticality: applying the criticality of a business case to a specific test case is where he says a human still has to sit, and he says test case creation is not a place to remove one.
  • Should agents run mission-critical systems autonomously? - No, according to Jeba Abraham, who answers "absolutely not" and adds that even if the technology reached that level it is not something Capgemini would recommend. Raghavendra Prasad MG sees the industry moving through human-on-the-loop to human-out-of-the-loop over five to ten years, and is explicit that this is a direction of travel rather than present-day practice.
  • What is a guardian agent? - Raghavendra Prasad MG runs one in parallel with every agent his team builds, at one guardian per agent rather than one per team or pod. It surfaces hallucination, toxicity and relevancy as a continuously running score so a human can act on it, and he says the pattern is implemented across multiple engagements.
  • Who watches the LLM that judges the other LLM? - Nobody yet, and Raghavendra Prasad MG says so plainly. Bringing in a second model to monitor the first is the LLM-as-a-judge pattern, and what happens when the judge is wrong is a question he says he does not have an answer to today.
  • How do you test a non-deterministic agent? - Jeba Abraham lists the dimensions Capgemini now covers that did not exist before: hallucination and drift, privacy exposure including PII and PHI, security including prompt injection, reliability measured as latency variance across contexts, sustainability, and demographic bias. The framework underneath is TMap risk-based testing, which is Capgemini’s own methodology.
  • What drives token costs up in agentic QE? - Uncontrolled uploads. Raghavendra Prasad MG describes engineers pushing whole business requirement documents and entire exported regression suites into prompts, and compares the metered bill that follows to an electricity bill landing on the organisation. His fix is a layered knowledge or context fabric so an engineer uploads only a user story.
  • What does it take to get an agent approved inside a bank? - Raghavendra Prasad MG says at least 25 people across the account, with the CFO and the security officer involved, full source code disclosure, and a connection over the client’s own network. Ninad P Shirodkar adds that BFSI clients rarely accept external products at all, so agents have to be built inside the client environment.
  • Is observability now part of quality engineering? - Yes, and Jeba Abraham treats it as a structural change rather than a tooling preference. Tools normally associated with AI operations have moved into the testing conversation, because validating an agent once and putting it into production without watching it is the failure mode to avoid.
  • What KPIs should QE leaders track now? - Jeba Abraham says retire test cases executed and automation coverage in favour of release confidence, customer experience and risk reduction. The reasoning is that metrics pull everything behind them, so team structures, tooling investment, adoption and skills all follow once the scoreboard changes.

The Broken Headcount Equation

Khimanand Upreti answers first and organises the change into what he calls three W’s: work, worker and workplace. He then picks one, saying the discussion should focus on work, and he never returns to the other two.

His core claim is that AI supplies scale without proportional hiring. Headcount may still have to rise in some situations, he says, but the rate of increase will not be what it was.

The role shift he describes is from writing test cases to supervising output, which he says raises rather than lowers the demand for human judgment, because somebody has to review what the machine produced.

He turns that into a measurement problem the panel keeps returning to. If a good tester is no longer the person writing the most test cases or finding the most defects, and the definition becomes something like better human judgment, nobody yet knows what the KPI for that looks like.

The New Engineering Pyramid

Ninad P Shirodkar contrasts the old structure of developers, testers, test analyst, test lead and test manager with a three-layer arrangement he says is forming now.

At the base are engineers and developers who use AI in day-to-day work. In the middle are SDETs and AI engineers who build the platforms and the agents themselves. At the top are AI architects who compose and orchestrate agents and handle context setting.

He treats the middle layer as the genuinely new one. Building platforms and agents requires a different type of skill, in his framing, rather than existing automation work under a new name.

His argument about foundations is the sharper point. A computer science graduate once learned fundamentals and could then pick up Java or .NET or whatever the client used. The equivalent foundation now, he says, is prompt engineering, the ability to create knowledge graphs, and the ability to create MCP servers.

His practical advice is depth before breadth. Become expert in one AI tool and you can pick up whichever tool the client happens to run, exactly as engineers used to move between languages.

He justifies the urgency with tool turnover, describing a flood of AI tools arriving over the last few quarters. He cites Cursor as his example and stumbles over both its launch year and a large dollar figure attached to it, neither of which is recoverable from the recording, so neither is reproduced here.

Dev-To-QA Ratio Reset

Jeba Abraham gives four trends. The first is that the dev-to-QA ratio has changed significantly, followed immediately by a refusal to quantify it, which is worth respecting rather than filling in.

The second is the question clients and teams now put on the table. Why do we need testers at all, when there are agents, when the code is being vibe coded, when the agentic framework already ships a testing persona that will write the Playwright test? Jeba Abraham is voicing that question rather than endorsing it, and answers it in the negative twenty minutes later.

The counterweight offered is context. Not all applications are built equally, and testing a mobile app is not testing a core banking system or something running operations for a large manufacturer, so team composition has to follow the difference.

The third trend is that automation has stopped being a downstream afterthought, and the conversation has moved past automating tests to autonomous workflows spanning requirement selection through to defect logging.

The fourth is that quality engineering is becoming more platform driven than people driven, with the caveat that when everything can be integrated with everything else through MCP, deciding what is worth integrating becomes the leadership job.

The most quotable of the four is a lot less gray hair. Jeba Abraham sees younger, near-AI-native engineers arriving on engagements, on the reasoning that experience brings baggage, and habits have to be unlearned before new ways of working can be built in.

Note

Note: Prove what the agent did, not just that it ran. Try TestMu AI now!

The Intent-Based Vision

Raghavendra Prasad MG argues by analogy. People who first got calculators redid every calculation on paper to check the machine, and only trusted it blindly after a couple of decades. The dates he attaches are shaky and are left out here; the shape of the story is the part he is using.

His second analogy is accounting software. Finance staff ran the package alongside a paper ledger for years, and eventually the ledger went away.

From those he projects toward what he calls intent-based implementation, where what you want is the only thing that matters because it aligns directly to business intent, and everything beneath it is assured as a commodity.

The trajectory he sketches runs from human in the loop today, to human on the loop once orchestration matures, to humans eventually coming out of the loop over five to ten years.

He frames the commercial consequence bluntly: this stops being a mass business and becomes a niche, class business, and he says that is what his large customers are asking for.

When the host pushes back on pushing the human out, he clarifies rather than retreats. It is a vision and it is not there today, he says, and then restates the direction of travel through the same calculator analogy.

Quality As A Platform

Asked how quality-as-a-platform lands in practice, Jeba Abraham says partnerships have become far more critical, and that a lot of leadership time now goes into talking with partners.

The distinction drawn is between partnering and selling. Working together on proposals still happens, but the newer part is co-innovation, with interoperability between platforms as the practical problem.

The worked example is integration mechanics: how to bring Playwright in as part of HyperExecute, and what to do when another tool has to sit alongside it. HyperExecute is a product of the company publishing this recap, though the framing in the session is generic interoperability rather than a recommendation.

Internal standardisation is the second shift described. Centres of excellence and accelerators have always existed, and the newer work is getting different groups inside one organisation to converge on a single point of view about what quality engineering now means.

The stated goal of that consolidation is external as much as internal, in the form of consistent messaging to the market. Jeba Abraham explicitly parks the people half of the answer for the later reskilling question.

The Agentic Split

Ninad P Shirodkar says his organisation split the problem into an engineering part and a business part, and that the ambition is no longer productivity but hyperproductivity. No before-and-after numbers accompany that claim.

He describes a QE platform his organisation built that maps the test lifecycle into three zones, requirements, test execution and continuous improvement, and then sorts each activity into fully agentic or human-in-the-loop.

Fully agentic in his mapHuman in the loop in his map
Impact analysis based on requirementsTest data condition identification
Test case and test scenario creation, which he calls his favouriteAll reviews
UI automation script generationTest data management
API automation script generationPairing and data-driven automation work, described in terms the captions render unclearly

Test script migration is the one activity he names under continuous improvement, and he places it in neither column, so it is left unsorted here rather than assigned.

The deployment constraint he names is specific to banking and financial services. Clients build their own frameworks and will not take external products into their environment, so agents have to be built inside the client’s environment, and built quickly.

His conclusion is a selection discipline rather than a coverage one. Identify the use cases that deliver the maximum benefit and build the framework around those, instead of attempting everything at once.

Validating Autonomous Agents

Raghavendra Prasad MG concedes the human-in-the-loop point first, on maturity grounds, then names agents where his teams have already removed the human. He offers them as examples rather than a closed list: requirement quality, scenario identification, synthetic data generation and self-healing.

He caps that immediately. The autonomy is a calculated risk taken where cost reduction can be demonstrated, and test case creation is not somewhere he removes the human.

His account of BFSI approval is the most concrete thing in the session. Putting one agent inside a bank means talking to at least 25 people in the account, the CFO and the security officer get involved, all the Python source has to be shown, and the agent connects over the client’s own network.

His monitoring analogy is a certified driver. The person passed the test, and you still watch for speeding in the city or harsh braking, with an alert and a scoring report following. That is how he frames continuous agent observability.

He names an eval stack used on retail and insurance engagements, plus an in-house solution built on custom embeddings. Ragas is clearly audible, as is a paid tool he calls Confident; the remaining product names in that list are garbled in the captions and are not reproduced here rather than guessed at.

Then he states the unsolved recursion openly. Bringing in a second model to judge the first is the LLM-as-a-judge pattern, and what happens if the judge is wrong is a question he says he has no answer to today.

Testing Nondeterministic Output

The host re-asks the autonomy question after Jeba Abraham rejoins the call. The answer that follows is not a response to Raghavendra Prasad MG, who had answered the same prompt during the disconnection.

Comma

That is reconciled with the earlier ratio point rather than contradicting it. There may be a lot fewer humans than before, and humans are still needed to validate certain things, provide judgment and make the critical calls.

For method, Jeba Abraham points to TMap, the test management approach, described as around 30 years old and as a framework bible for risk-based testing, with a release last year covering validation of non-deterministic output. TMap is Capgemini’s own methodology, and the authorship is claimed on air.

The list of what now gets tested is the useful part: hallucination and drift on the functional side, privacy in terms of PII and PHI exposure, security including prompt injection, reliability measured as latency variance and behaviour across different contexts and environments, sustainability, and demographic bias.

The structural observation is that observability as a discipline has moved much closer to quality engineering, with tooling normally associated with AI operations now needed to test agents. Of the three tools named, only Langfuse is clearly audible in the recording.

The operating model described is crawl, walk, run. Build the agent in a couple of weeks or considerably less, harden it, then scale, with token economics tracked as it scales, and an explicit warning against validating once, shipping, and forgetting about it.

Shift from a legacy test platform to TestMu AI

Context Fabric And Token Cost

Raghavendra Prasad MG frames ungrounded agents as a compliance hazard, using an example he flags as hypothetical: an engineer in India asking about an Australian bank, where the agent could get confused and apply US regulation.

His fix is what he calls a knowledge fabric or context fabric, which he says solves two problems at once by making the agent domain-aware and cutting token spend.

The token problem he describes is behavioural rather than technical. When engineers are enabled individually at project level they upload business requirement documents, requirement documents, and entire exported regression suites, and the metered bill that follows arrives at the organisation like an electricity bill.

He identifies the ownership problem underneath. The cost does not sit with the QE organisation, it sits at organisation level, so when QE proposes a retrieval layer the sponsor often turns out not to be the owner, which is why maintaining that layer is itself a challenge.

His layered design has three tiers. A business layer that changes rarely, with a central bank KYC rule change as his example. An application layer that moves only on modernisation, such as Java to Angular or a packaged platform to a custom build. A transactional layer carrying agile artefacts, tickets, defects and production defects.

The payoff he claims is that an engineer uploads only the user story and retrieval supplies the rest. He also insists the output stays probabilistic, joking that an agent asked the same question repeatedly will get annoyed like a person would.

Guardian Agents And Guardrails

Raghavendra Prasad MG says the eval space now offers roughly 200 metrics, covering hallucination, F1, precision, recall, accuracy and deviation, and that his teams select only a small subset of the top ones to measure a given agent. The exact fraction he gives is garbled in the captions and is not reproduced here.

Comma

When the host restates it as a parallel auditor per team pod, he corrects the granularity. It runs per agent, not per pod and not per layer.

The guardian surfaces industry-standard metrics as a running score, and the three he names are hallucination, toxicity and relevancy, so a human can act on the number rather than inspect the agent.

Khimanand Upreti reframes token optimisation as a consulting opening, putting it conditionally: advising customers on optimising token cost would get a partner into the ecosystem easily, because not every enterprise uses tokens optimally.

He notes vendors shipping control towers and agentic control planes at varying depths, some with three layers of security and some with four or six, all aiming to keep an agent doing only what it is supposed to do. Alongside that he mentions a platform like TestMu AI as useful, flagging on air that he is saying so at that company’s own conference, and framing it as an example of a consolidating enterprise platform rather than a specific recommendation.

He then punctures his own confidence with a news story he had recently read, about an agent that deleted a customer’s production data completely. No company, date or publication is named, and his conclusion is that he could not put a finger on total certainty that an agent will do what it is supposed to do.

A short exchange follows on ungoverned agents, excessive access permissions, identifying non-human agents, and every agent carrying an attack surface. The captions show at least two voices inside twenty seconds with no name attached to any of them, so none of it is attributed here.

Honest Career Paths

Khimanand Upreti opens by citing a recent World Economic Forum report that he says predicts quality engineering is a discipline whose business will grow rather than shrink. No title, year or figure is given, and he hedges the citation himself.

He then gives three levels he sees in his own teams: engineers using AI tools for productivity, engineers who go beyond using AI to actually build agents and models, and AI enterprise architects advising on the whole ecosystem of processes, governance, tooling, skilling, resourcing and environments.

His point about the middle level is the interesting one. Building the agents needs people who carry a quality engineering mindset, rather than the mindset being something bolted on afterwards.

Ninad P Shirodkar lays out a ladder of four named levels: engineers and developers using AI, SDETs becoming AI test engineers who build agents and workbenches, forward deployed engineers who productionise those toolsets alongside AI architects, and at the top agent orchestration leaders who orchestrate platforms and build business outcomes.

He pairs that with three skilling levels his employees go through: an AI foundation level covering GenAI basics, prompt engineering and one toolset, an AI-assisted level covering agent creation, retrieval and IDE-driven generation, and an advanced level for architects and orchestrators. Only the foundation level is described as compulsory for every employee.

Jeba Abraham names three honest paths. Deeper domain knowledge, so someone can judge business intent, citing the Australian banking regulation example directly. The trust layer and governance, covering observability and the testing of agents, described as work being done now that was not being done a year ago. And generalist engineering, where the specialised automation engineer and the performance engineer converge and SDETs become forward deployed engineers, closer to a Swiss army knife.

Raghavendra Prasad MG gives a target role mix his employer is working toward: 10 to 15 percent of engineers upskilled into a GenAI specialist role, 20 percent into forward deployed engineer roles, and 60 percent into AI engineer roles, meaning the engineers who use the agents rather than build them. The three figures are an aspiration rather than a completed reorganisation, and they do not sum to 100.

He adds a fourth, deliberately scarce role, the governance specialist, at perhaps one or two people for the largest engagement, running a marketplace covering agent orchestration, governance, token economics, estimation and pricing agents. He credits the idea to Khimanand Upreti, says the marketplace is still in progress rather than delivered, and is explicit that such people cannot be produced like an army.

Three-Year Outlook

Khimanand Upreti’s forecast is a word order swap he says carries real weight: it is human plus AI today and it may become AI plus human. He also predicts testing stops being a downstream function, and that the boundaries between manual, automation, functional and performance testers, and eventually between dev and QA, will diminish.

Raghavendra Prasad MG argues quality engineering survives but shrinks into an expensive niche. His analogy is an excavator: fifty people once dug a hole, and now one operator does it in an hour and charges three thousand rupees for that hour, which makes it an expensive skill rather than a mass business.

He then goes on the record about his employer’s commercial posture on large accounts of 300, 600 and 700 people, saying they now go in up front offering to cut people cost by 50 percent, alongside agent-based pricing so that the agent eventually earns the money. That is a forward-looking offer being made to clients, not a delivered result, and he presents it as his own account of how his firm is bidding.

Jeba Abraham tells leaders to change the scoreboard first. Move the KPIs off how many test cases were executed and what the automation coverage is, and onto release confidence, customer experience and risk reduction.

The reasoning is that metrics pull everything behind them. Shift the metrics and team structures change, tooling investments become clearer, and adoption and skills development follow.

Ninad P Shirodkar’s answer is cultural and top-down. It begins from the top, he says, citing that his company’s chairman personally runs masterclasses on the subject, and he wants every leader hands-on: take an AI tool, start building agents, start building workbenches. The host closes by predicting that after prompt engineering the role becomes trust engineering, and that quality engineers are the closest fit to be its gatekeepers.

This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.

Author

...

TestMu AI

Blogs: 232

  • Twitter
  • Linkedin

TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests