Next-Gen App & Browser Testing Cloud
Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Engineering Trust Quality in AI-Driven Commerce [Testμ 2026]
Engineering Trust Quality in AI-Driven Commerce [Testμ 2026]
A Testμ 2026 retail panel on continuous readiness, bot traffic already being blocked, and why agents test the data layer rather than the storefront.
Published on:
On This Page
- Super Bowl Sunday At Pizza Hut
- Black Friday To Christmas
- New Look's Pre-Peak Checklist
- Many Peaks, Three-Legged Race
- A Business Built On Weather
- Continuous Readiness
- Bots Already Being Blocked
- The End Of The War Room
- The 30 To 40% Figure
- Into The Data Layer
- Trust Deep In The Platform
- Testing Without One Answer
- Five Closing Shifts
- Q & A Session
The retail quality playbook has a date on it. Load-test the funnel in August, freeze the code in October, staff a war room in November, and find out on one graded day whether the year’s work held.
In this panel discussion at Testμ Conf 2026, five quality leaders answered from five retail seats: Sunita McCoy, Director of Quality Engineering at Pizza Hut; Rama Gunjaipalli, who leads quality engineering at Michael Kors; Krishna Gangu, Distinguished Architect at Catalyst Brands; Nipam Desai, Engineering Manager for Quality, Environment and Release at New Look; and Jason Bryant, Director and Head of Software Quality Engineering and Assurance at Tractor Supply.
They still have that graded day. What they no longer have is confidence that preparing for it as a date works, because a growing share of the traffic arriving at their sites is not a person, and the same agent can run the same scenario two different ways.
If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.
TL;DR
Agentic commerce is retail in which a growing share of site traffic is AI agents that browse, compare and increasingly transact on a shopper’s behalf. Five retail quality leaders took it up at Testμ Conf 2026 because it moves the object of testing off the rendered storefront and down into the data layer, the APIs and the contracts.
- Does readiness still mean green dashboards? - No. Sunita McCoy of Pizza Hut defines readiness as systems holding under heat with nobody white-knuckling, and success as a quiet, incident-free Super Bowl Sunday that is boring for her team. Green dashboards are part of it, not the whole of it.
- Is the 30 to 40% bot-conversion figure Pizza Hut data? - No. Sunita McCoy of Pizza Hut attributes the figure to data and metrics that were shared, without naming a source, and says it appeared close to that range rather than asserting it. No sample, baseline, method or definition of channel conversion accompanies it, and the published chapter list presents it far more confidently than she did.
- Are retailers letting agents transact freely today? - No. Rama Gunjaipalli, who leads quality engineering at Michael Kors, says his organisation is not extensively letting bots crawl its systems yet because it wants to understand their intent first, and answers his own question about allowing all bots to transact with a not yet, followed by yes, that is the direction.
- Has bot blocking already caused problems? - Yes. Jason Bryant of Tractor Supply had one of his own agentic test platforms blocked while it ran post-production validation. He does not name the platform, and he sets the incident against the opposite risk of blocking a genuine agent-driven sale.
- What does an agent actually test against? - Sunita McCoy of Pizza Hut and Krishna Gangu of Catalyst Brands both answer the data layer, the APIs and the contracts rather than the rendered storefront. Bots and agents are not really looking at your store, in her framing, and agents reach the platform through APIs, in his, so the storefront alone is no longer the thing under test.
- Where does trust get enforced now? - Deep in the platform rather than at the edge, in Krishna Gangu’s formulation, through data governance, knowledge graphs, federated per-domain ownership, data lineage and canonical models, so price, inventory and promotion match whether the answer surfaces on your own site or through an external assistant.
- Does given-when-then still work for agentic features? - No. Nipam Desai of the fashion retailer New Look says the pattern breaks down when several answers are valid, so acceptance criteria become criteria boundaries: check the item is genuinely in stock, and check the price boundary accounts for promotions and flash discounts, rather than checking for one expected output.
- What replaced the peak war room at Michael Kors? - Thorough validation in lower environments built as close to a production replica as possible. Rama Gunjaipalli describes the war room era as intense, staffed around the clock through the holiday period, and says he would rather put the effort into testing than find issues in production. The evidence offered for the change working is subjective confidence, not a defect-escape rate.
- How does an agent earn trusted access? - Jason Bryant of Tractor Supply flags his own answer as circular before supplying real criteria: accuracy, not going rogue, transparency about what it is doing, consistency across runs, and recovering quickly when behaviour turns strange.
- Can hallucinations be caught automatically today? - No. Sunita McCoy says the only way she can confirm something is hallucinating is a human in the loop, which makes a subject matter expert with years on the ground before AI her most valuable organisational asset. She adds outright that she has no golden solution yet.
- What are the three legs of readiness? - Platform readiness, operational readiness and business readiness, in Krishna Gangu’s framing. Operational readiness is runbooks, monitoring and proactive alerting rather than code, and business readiness is a set of merchandising questions about having the right content, prices, inventory and products available on time.
- Were audience questions answered in this session? - No. The moderator noted a full Q&A queue, skipped some prepared questions to reach one more panel topic, and twice apologised on air that there was no time for audience questions. Not one was read out.
Super Bowl Sunday At Pizza Hut
Sunita McCoy names her graded day as the most important day for pizza in North America.
Her figures are self-reported with no source, sample or baseline behind them: over 1.5 million pizzas that evening, a peak at 6pm Eastern where the chain sells about 5,000 pizzas a minute coast to coast, and more than half of that arriving through the digital channel.
The first two at least corroborate each other. Five thousand a minute sustained across roughly five hours lands near 1.5 million, which makes them internally consistent even though neither is externally verifiable.
Her definition of ready is the part worth keeping, and it is deliberately not a metric. Readiness does not restrict itself to green dashboards; the systems have to hold under heat and nobody should be white-knuckling.
Success, in her framing, is the absence of drama. A quiet, incident-free Super Bowl Sunday is boring for her team, and boring is the goal.
Black Friday To Christmas
Rama Gunjaipalli names Black Friday for what he calls obvious reasons: as a fashion retailer, the majority of the sale is expected during that period. No figure is attached to the majority.
His readiness bar is zero tolerance, on the reasoning that customers are impatient and cannot be shown any delay on the website.
Nipam Desai extends the day into a window. For New Look, peak starts at Black Friday and runs to Christmas, with pre-peak trading preparation already under way at the time of recording.
His reason the period functions as the exam is that the volume of customers and the volume of transactions peak together, so the whole year of quality work and every feature released into it gets tested at once.
He is the only panelist to follow the transaction past checkout, noting that an order placed on a digital channel or marketplace transfers to operational and back-office systems, and that those downstream systems get validated at peak too.
His summary of the mood is a hedge rather than a boast: the digital engineering team is always crossing its fingers, ready but hoping for the best.
Jason Bryant supplies the one piece of audience level-setting, explaining that US retailers call the day after Thanksgiving Black Friday because the year’s books move from red into black, and adding that Tractor Supply internally just calls it the day after Thanksgiving.
New Look’s Pre-Peak Checklist
Nipam Desai lists four preparation workstreams: the right level of automation and regression coverage, every critical customer journey covered and validated end to end, extensive performance testing, and extensive resilience testing.
The goal he states for the package is visibility rather than a clean bill of health. He wants clear sight of the remaining risk, and a mitigation plan for anything that gets in the way.
The most concrete item on the list is also the least glamorous. It is a contact list: who the key point of contact is in each department and team, and how to reach a third party if something goes wrong.
None of this is offered as new. It is the traditional playbook, and it is the baseline against which the panel’s later answers about what no longer holds should be read.
Nothing in the checklist was shown. There is no dashboard, no coverage report and no test-suite screenshot anywhere in the session.
Note: When the buyer is an API client, the storefront stops being the system under test. Try TestMu AI now!
Many Peaks, Three-Legged Race
Krishna Gangu reframes the question rather than answering it. A multi-brand consortium has many peaks, and he names Memorial Day, Labor Day and back-to-school alongside the Black Friday to Cyber Monday window, while conceding the fourth quarter is critical for retailers.
He describes Catalyst Brands himself as a portfolio holding JC Penney, Brooks Brothers, Aeropostale, Nautica and Lucky under one private equity owner, formed the year before the recording.
His account of the traditional cycle is precise and he says it still partly persists: preparation from around August through October, then code freezes and testing the platforms under real load. His stated purpose for it is commercial, so that customers are not embarrassed.
The break comes from architecture. Cloud-native platforms with multi-tenancy and a cross-region presence push companies toward continuous failover testing, which he states as a prediction about where the industry is going rather than as something already finished at his own employer.
His organising frame is a three-legged race between platform readiness, operational readiness and business readiness. Operational readiness is runbooks, monitoring, proactive alerting and resilience, meaning the human and process layer rather than code.
Business readiness he expresses entirely as merchandising questions: whether there is the right content, the right prices, the right inventory available on time, and whether the right products are shown to a customer who is searching. Platform readiness is the leg he says is changing, toward continuous failover and continuous performance testing.
The moderator then reuses that three-part framing almost word for word to set up her next question, so it propagates through the rest of the hour as host scaffolding rather than as agreed panel consensus.
A Business Built On Weather
Jason Bryant is the one panelist whose peak has no date. Tractor Supply is cyclical, and the business is built on the weather being outside.
He defines the customer base by activity rather than demographics: a lifestyle and culture of people farming, caring for animals outside and raising horses.
The trigger is seasonal rather than calendared. As soon as it warms up and flowers bloom, business picks up, so the team has to be ready the moment spring hits and hopes it hits early.
He ties that to reported results, saying an early spring helps make the first quarter’s earnings in that March month and shapes how the year goes. The claim is directional, with no figure behind it.
The conventional peaks still exist on top of that baseline, with the day after Thanksgiving his biggest day and Cyber Monday another. His closing wish is the clearest statement of why his readiness cannot be a single event: he wants it warm, and warm all year.
Continuous Readiness
Asked which pre-agentic assumption no longer holds, Jason Bryant answers flatly that readiness is not a one-time thing anymore. It does not mean getting ready once and declaring yourself ready; it means consistently getting ready.
Continuous readiness is his company’s own name for the practice, and preparation starts right before the summer months to cover a spring-triggered business.
What it consists of is process rather than tooling: cross-functional meetings, monitoring set up correctly, current vendor contacts, and internal and external partners ready to jump if something goes wrong.
He characterises the change methodologically. Getting ready is no longer a waterfall approach but an agile one, done continuously and iteratively.
The agentic reason he gives is unpredictability, and he states it twice: nobody can predict what the bots will do, and the same bot could run the same scenario one way this time and differently the next.
The honest limit of his answer is that his big focus remains operational stability and performance testing. The practices are traditional; the cadence is what changed.
Bots Already Being Blocked
Pressed on whether he has seen the bot activity already, Jason Bryant hedges rather than confirming outright. He thinks so, and knows his team is constantly blocking new things interacting with the site.
The session’s best first-hand incident is his own tooling getting caught in that net. One of his agentic test platforms was blocked while trying to run post-production validation testing. He does not name the platform, and because the publisher of this recap sells agentic testing tooling, no inference about which one is made here.
He states the counter-pressure using named third parties, raising the case of a shopper going to Perplexity and buying from Walmart through its integration. That integration claim is his, second-hand and undated.
He locates the problem organisationally rather than technically, saying his information security department is really having to juggle right now.
Rama Gunjaipalli’s parallel position carries a qualifier the chapter list flattens. He admits his organisation is not extensively letting these bots crawl into its systems yet, because it wants to know the bots’ intentions first. The word extensively is in the recording and changes the claim.
He names the real difficulty as intent classification: telling apart a bot probing for security vulnerabilities, a bot crawling, and a bot transacting the way a customer would.
The End Of The War Room
Rama Gunjaipalli describes the era he is glad to have left. Peak readiness meant a war room with every vendor on standby in case of an issue, and he calls it quite intense.
He is specific about the human cost rather than dressing it up. Teams worked shifts around the clock through the holiday period, which he says is not great, and is real pressure to be present just in case.
What replaced it is environment fidelity. Most validation now happens in lower environments built as close to a production replica as possible, so issues that would have surfaced in production have already been handled.
He self-corrects mid-sentence on how much testing that means, starting to say excessive and landing on thorough due diligence. The correction is preserved here because smoothing it changes the claim.
His preference is unambiguous. He would never want to be in a war room situation, and would rather put more effort into testing than find issues in production later.
The evidence offered that this works is entirely subjective: a great deal of confidence for the business and IT teams, and his own satisfaction with the shift. No defect-escape rate, incident count or before-and-after accompanies it.
The 30 To 40% Figure
The session’s most quotable statistic is also its weakest. Sunita McCoy says that in the last holiday season, from data and metrics that were shared, it appeared close to 30 to 40% channel conversion through bot traffic.
Every hedge in that sentence is hers. Shared by whom is never said, she says it appeared close to rather than asserting it, and she does not present it as Pizza Hut’s own data. There is no sample, baseline, method or definition of channel conversion attached.
The published chapter list converts all of that into a flat statement of last season’s data, which is the single most misleading piece of metadata attached to this video.
Her conclusion is the part that survives the number: if you are not accounting for agent traffic, you are leaving money on the table. That argument holds whether or not the figure does.
She sets it up with the panel’s cleanest picture of the strategic reversal. About twelve months earlier the position was to restrict everything with web application firewalls, CAPTCHAs and rate limits. Now the conversation is about opening up.
She frames what follows as open questions rather than answers: whether an agent placed a true order, whether the card is real, and whether that revenue is genuinely being allowed into the organisation.
Into The Data Layer
Sunita McCoy states the shift most compactly. Bots and agents are not really looking at your store, so what needs validating is the underlying data layer and the stability of the API calls.
Her conclusion is that testing moves away from what meets the human eye and toward contract-based testing.
Krishna Gangu makes the same move in platform terms, correcting himself as he goes: agents access the APIs and the platform, not the sites, so what is under test is the platform with its APIs rather than the storefront alone.
Rama Gunjaipalli’s version is about scope rather than layer. His plan is to test not only the interface a human customer sees but to factor in the bots as well. The captions drop a word from that sentence and reverse its sense, so it is paraphrased here and not quoted.
Nipam Desai adds the consumer and provider dimension, describing extensive contract testing that validates both sides thoroughly, including control over backdated versions.
Jason Bryant’s version is about granularity. With decision-making bots it is less about one piece of functionality working and more about the whole decision-making process, which pushes testing toward an end-to-end perspective.
Trust Deep In The Platform
The mechanisms he names for that are data governance, knowledge graphs, federated per-domain data ownership, data lineage and canonical models.
The consistency requirement he sets is concrete enough to test. Whether a customer arrives through an external assistant, the mobile app or the site, what gets served is a relevant product with the right price, the right inventory and the right promotion at any given moment.
His argument about discovery follows from that. Crawling and clicking give way to answerable content rather than clickable content, so the test question becomes whether content is grounded in enterprise data and surfaced correctly across multiple answer surfaces.
Jason Bryant’s definition of when an agent earns access is circular, and he says so himself before supplying the real criteria: accuracy, not going rogue, transparency, consistency, and recovering quickly from strange behaviour.
Krishna Gangu also offers a second-hand anecdote about a restaurant chain whose public conversational chat was scripted against by the public, which he says was big news. He gives no date, no outcome and no publication, and it is not corroborated anywhere in the session, so it is reported here as his recollection rather than as an established incident.
Nipam Desai frames the whole shift as upside rather than risk, calling agentic commerce a completely new revenue stream and a successor to social commerce. His accountability standard is explainability: if an agent buys on a customer’s behalf, the team must be able to explain the decision and show evidence it was right.
Testing Without One Answer
Nipam Desai names the break precisely. The traditional given-when-then criteria expect a certain output from a certain input, and in an AI era there are several valid answers, so the question becomes what criteria boundary to draw.
His worked example is a search feature his team is building. A shopper asks for a summer item under a stated amount, and the output could be any of ten different things. The garment word is an unclear rendering in the captions, so the example is described rather than quoted.
Validation then becomes a set of factors. The item has to actually be in stock, because there is no point in an assistant returning something unavailable, and the price boundary has to account for promotions and flash discounts that pull an item into range.
Sunita McCoy proposes an explicitly hybrid strategy and admits she has not solved it. There is a deterministic set of validations for the golden customer journeys governed by business rules and engines, paired with non-deterministic validation for everything else.
Her hallucination check is a person, stated in the first person. The only way she can confirm something is hallucinating today is a human in the loop, which makes her most valuable organisational asset a subject matter expert who was on the ground long before AI arrived. The published chapter title turns that personal statement into a universal claim she did not make.
Her aphorism for the shift is that cash is no longer king and context is king. The sentence is disfluent as spoken, and the chapter list puts a tidied version inside quotation marks.
Five Closing Shifts
The moderator asks each panelist for one shift, and gives them roughly a minute each.
- Nipam Desai, New Look: move from testing the journey to testing the decision the agent is allowed to make.
- Krishna Gangu, Catalyst Brands: get the data strategy knocked out first, because the AI is only as good as the data, then answer-surface optimisation, then federated canonical catalogue feeds, then agentic checkout.
- Rama Gunjaipalli, Michael Kors: equip the security team with the tools to judge what to allow, and make the site as bot-friendly as possible for transactional purposes.
- Jason Bryant, Tractor Supply: keep it simple. A bot that handles returns should not try to handle every possible type of return, and teams should not try to boil the ocean.
- Sunita McCoy, Pizza Hut: start with people rather than technology, given three or four generations at work right now, then build a measurable playbook covering platform access, adoption rate and productivity gain, and use one solid end-to-end journey to build internal trust.
Q & A Session
There was no audience Q&A. The moderator noted a queue of questions and skipped a couple of prepared items to reach one more panel topic.
Two exchanges in the hour do function as questions and answers, both put by the moderator rather than the audience.
- Have you already seen that sort of bot activity and learned from it?
Jason Bryant: I think so, and I know my team is constantly blocking new things interacting with the site. An agentic test platform of my own was blocked while running post-production validation. A shopper can go to Perplexity and buy from Walmart through its integration, and my information security department is having to juggle. A hedged yes. The blocked platform is the session’s one first-hand incident and he never names it, and he raises the opposite risk immediately.
- Is Black Friday as big a peak in the UK?
Nipam Desai: It is becoming big and getting a lot more traction, and anything celebrated in the US eventually gets adopted. A directional answer with no figures.
This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.
Author
TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




