Hero Background

Next-Gen App & Browser Testing Cloud

Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

Next-Gen App & Browser Testing Cloud
Testμ

Building a Billion Dollar Healthcare Product at Zebra Technologies [Testμ 2026]

Zebra put QA at the front of product discovery - the nurse documentation problem, the pivot to dictation, and the internal AI tools behind it.

Published on:

Nurses told Zebra’s team they were spending almost half the day documenting patient data, on an electronic transfer system that had not changed in years.

That finding arrived only because the team had already tested something else and been wrong. Their first bet was emotional support. Nurses redirected them to admin burden.

At Testμ Conf 2026, Spyros Katopodis, Senior QA Lead at Zebra Technologies, walked through that discovery loop and the internal tooling behind it.

Youtube thumbnail

If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.

TL;DR

Zebra’s AI QA stack is a set of internal tools that turn requirements documents into executable tests, paired with a shift-left process that tested product hypotheses on nurses before anything was built. It produced a wearable assistant that transcribes patient conversations and passes handoff tasks between shifts.

  • Does Zebra run QA before development rather than after? - Yes. Zebra Technologies places quality assurance at the front of product discovery, and the speaker opens on that contrast, saying many companies treat QA as the last check after development. His framing is that the unlock in finding the next product was rethinking how AI-powered QA could shape and scale the work from the start.
  • What problem statement did the healthcare experiment start from? - Making nursing so satisfying that nurses could not wait to go back to work. That was translated into four goals: help nurses work more efficiently and more effectively, give them emotional support, and increase job satisfaction. The emotional-support goal was tested first and dropped first.
  • Why was the emotional-support idea abandoned? - Zebra dropped the emotional-support hypothesis because nurses named something else. Asked directly, they identified time-consuming repetitive admin work, particularly filling out documentation by hand, as their main source of stress, which the speaker says made it obvious the need was automation rather than support.
  • How much faster was dictation than typing in Zebra’s comparison? - 50% faster, in a comparison the team ran on themselves rather than in a hospital. The speaker names accuracy problems alongside it, including misspellings and dropped content, and gives no sample size, participant count or measurement protocol.
  • How much time did nurses report saving per patient? - Nurses using Zebra’s speech-to-text approach reported saving four to ten minutes per patient, which is self-report gathered after the fact rather than a measurement Zebra took. The speaker chains it to several hours a day, several weeks a year, and several million dollars annually, without showing the arithmetic or naming a hospital size or patient volume.
  • How does the nurse-facing assistant work? - A nurse wears a Zebra device that listens to the patient conversation, transcribes the record and identifies handoff tasks. Updates reach the electronic health record only after the nurse approves them, and the tasks pass to the next nurse on shift.
  • Does the session address patient consent or data handling? - No. For a product that records patient conversations and writes to health records, the talk contains nothing on consent, de-identification, data residency, clinical validation or regulatory compliance. The nurse approval step is a clinician review of the record write, not patient consent.
  • What was the biggest bottleneck in the QA team? - Code refactoring, because of how long it took, followed by failure analysis, root cause analysis and debugging. Underneath all of it is the broader complaint that the team spent more time on admin, writing documentation and pulling status updates out of test plans, than on writing code.
  • What internal AI tools did the QA team build? - Eight are described: a feature-document generator reading Confluence, Jira and Figma; a mind-map and manual test case generator; a code generator producing tests for Cypress, Playwright and Robot Framework; an engineering insights dashboard; a no-code automation builder; a Robot Framework to Jira results API; a bug-ticket writer; and an in-house evaluator for AI responses.
  • How does the team turn failing code into readable bug tickets? - By sending the failing sections to a model. Converting automation code into human-readable reproduction steps was previously done by hand, line by line, so that whoever analyses the bug does not have to read test code, and their tool now generates both the ticket and the steps.
  • Are the reported time savings backed by data? - Not in this Testμ 2026 session. Zebra’s closing slide gives before-and-after ranges for test design, API test generation, root cause analysis and first pull request review, with no sample size, team size, measurement window, baseline methodology or date attached to any of them.
  • Was anything demonstrated live? - No. No tool was run during the Testμ 2026 session, and every one of Zebra’s internal tools was described verbally or shown as a static example slide. Two Zebra corporate advertisements were played in full, accounting for over a minute of the running time.

QA At The Front

Comma

His claim is that Zebra puts QA at the centre of what it does and at the forefront of everything. The exact noun in that sentence is garbled in the captions, so the phrasing here is a paraphrase.

He frames the whole talk as a behind-the-scenes account of an initiative that ran this year inside his organisation, aimed at uncovering the company’s next major business opportunity. He never states which year.

His formulation of the unlock is the thesis of the session. When the company went looking for its next product, the real unlock was not just the idea, it was rethinking how AI-powered quality assurance and testing could shape and scale that innovation from the very beginning.

The goal he names is turning testing into an engine for speed and quality instead of a gate at the end of the process. That phrase anchors everything that follows, and it is what separates this from a straightforward product case study.

From Cortexica To Zebra

He gives a compact career path: computer engineering and informatics in Greece, then an MBA, then a move to London for cloud security certifications focused on testing and monitoring.

His first job after that was at Cortexica Vision Systems, which he calls a very small but very mighty startup doing image search and visual recognition. Zebra took an interest in its technology and acquired the company, which is how he arrived.

He dates Zebra’s founding to 1969 and describes its purpose as helping companies monitor, anticipate and accelerate workflows by empowering frontline workers while keeping everything connected and visible.

He lays out the product arc as a sequence of expansions: barcodes and scanners, then mobile computing, then track-and-trace, then barcode printing, and only then software, built to integrate intelligence into the device. The installed base is the moat he names, calling the device fleet one of the company’s competitive advantages.

He then stops the talk to play a corporate marketing video. A frequently quoted line about decades of knowing the frontline comes from that advertisement’s voiceover rather than from him.

Defining The Frontline

He states that around 80% of the global workforce works in frontline operations, and offers no source, study or year for it. The claim also happens to size the market his employer sells into.

He defines the term carefully, because he leans on it: the frontline is everyone who provides the first human connection with an organisation, the public face representing the brand and the mission, interacting directly with customers and providing essential services.

The industries he lists are healthcare, transportation, logistics, warehousing, retail and manufacturing.

He gives a worked example per sector. In retail, insights that optimise inventory and staffing and anticipate customer needs. In logistics, data on routes, parcel dimensions and proof-of-delivery photos used to tune operations. In manufacturing, records of defects and equipment performance used to predict maintenance needs.

For healthcare he states the three-part goal that sets up the rest of the session: increase job satisfaction for healthcare professionals, increase efficiency, and increase hospital revenue. A second corporate advertisement plays shortly afterwards.

The Data Scarcity Claim

He describes the company’s AI portfolio as three layers: devices that support models, models that can see, hear, understand, think and interpret, and agents that distribute workload between device and cloud.

He then names what he calls the critical challenge facing AI, which is access to data. His claim is that the availability of high-quality, diverse datasets essential for training is becoming increasingly scarce.

His argument for why frontline data specifically matters is that it reflects in real time what is happening at the heart of a business, generated during daily operations and capturing the details of workflows, environments and interactions.

No source accompanies the scarcity claim, and it points directly at his employer’s commercial position as a frontline data-capture vendor. He does not say how the company obtains, licenses or anonymises frontline data, which is a notable gap given the healthcare use case that follows.

Note

Note: Test the hypothesis before you build the product. Try TestMu AI now!

The Twist Was Testing

He restates the mission and delivers the pivot of the talk. The twist is that it was not just about the data, and it was not just about the idea. It was about rethinking how QA and testing could shape and scale an innovation in the first place.

That word just carries the whole claim, and the published chapter list drops it, turning a both-and claim into an either-or one. He is not saying the idea did not matter, he is saying it was not the only thing that mattered.

His summary of the working method is that everything was about testing, pivoting and gaining results.

The two questions he says the team held while searching were how far AI could push their testing on efficiency and precision, and how fast it could move from an idea to real results.

The healthcare problem statement he read out is deliberately extreme: make nursing so satisfying that nurses and healthcare professionals could not wait to go back to work. He translates it into four goals, and the third of them, emotional support, is what the team tested first and abandoned.

The Nurse Feedback

He says the team connected with several experts, whom he calls hidden heroes, and established working relationships with several hospitals, gathering feedback from practising health professionals. No hospital, expert or nurse is named anywhere in the talk, and no count is given for any of them.

The headline finding is that nurses reported spending almost half the day documenting patient data.

The second finding is infrastructural. The electronic transfer document system was outdated and had remained unchanged for many years, which he frames as a lack of an efficient data transfer solution.

The team’s objective from that feedback was to give time back to nurses so they could focus on seeing patients.

He is explicit that the time saving was meant to convert into money, saying the team wanted to turn it into revenue and dollars for the hospitals, with more patients seen daily as the mechanism. The sentence in which he explains that mechanism is damaged in the captions, so only the intent is reported here.

The Pivot To Dictation

The first experiment was the emotional-support hypothesis. The team stepped out of its comfort zone, tested whether emotional support would increase job satisfaction, and verified it with healthcare professionals. It did not survive contact with users.

What nurses said instead was that their main source of stress was time-consuming, repetitive manual admin work, particularly filling out documentation by hand. He says it became clear and obvious that what they needed was technology and automation.

The team pivoted to testing whether entering patient data by speech would be faster, running a head-to-head comparison of dictation against manual entry into forms. He says they tested the solution themselves, so this was an internal comparison rather than a hospital trial.

The result he reports is that dictation was 50% faster, with accuracy problems including misspellings and dropped content. No sample size, participant count, form type or measurement protocol accompanies it.

Taking that back to the nurses, he says they reported saving four to ten minutes per patient, which he extrapolates to several hours a day for a nurse, several weeks a year for a hospital, and several million dollars annually. The per-patient figure is self-reported and the dollar figure is an extrapolation with no arithmetic shown.

The remaining experiments he lists are using AI to fix the accuracy and speed problems, testing natural language conversations on devices, evaluating shift-handoff challenges, and landing on task management plus ambient voice recognition, which he says nurses preferred.

Run tests up to 70% faster on the TestMu AI cloud grid

The Wearable Assistant

The solution he describes is a nurse-focused product providing shift handoffs and smart charting. The captions garble its name beyond recovery and the published description does not name it either, so no product name is printed here.

The workflow he walks through: a nurse wears a Zebra device, the device listens to the conversation, collects the data, transcribes the patient record, and identifies tasks for shift handoff.

The human step is explicit. Updates reach the electronic health record only once approved by the nurse, and the identified tasks then pass to the next nurse on shift.

He names Zebra Companion as a broader assistant you can talk to and ask questions, listing ambient voice recognition, task management, natural language understanding, on-device translation, in-moment learning, and recognition of text and barcodes. He introduces it separately from the nurse solution and never states the relationship between the two, so they are not treated as the same product here.

He adds that it can be customised for any organisation because it is trained on any company’s data, which he says ensures relevant and accurate responses. That is an absolute accuracy claim and an unexplained data-handling claim, and it is reported as his assertion.

One absence is worth stating plainly. For a product that records patient conversations and writes to health records, the session contains nothing about patient consent, de-identification, data residency, where transcription is processed, clinical validation or regulatory compliance. The nurse approval step covers the record write, not the recording of the patient.

The QA Admin Burden

He names the core problem in his own teams plainly: they were spending a lot of time on admin rather than actual code generation, specifically writing documentation by hand and gathering information out of test cases and plans to produce status updates.

The bottlenecks he ranks are code refactoring first, because of how long it took, then failure analysis, then root cause analysis, then debugging.

Where he says AI moved the needle is automatically generated test cases, execution with maintenance costs kept low, faster automation with self-healing, and no-code automation letting people create automated tests without knowing how to code.

He mentions a shared-tooling practice that is easy to miss: the team uses agents and skills shared between team members and across teams. He credits a terminal coding assistant alongside the internal tools for faster review, simpler failure analysis and faster debugging, though the product name is not cleanly audible.

Two further capability claims sit in this stretch and are absent from the published chapter list. He says AI helps predict which parts of an application are prone to defects so high-risk areas can be prioritised, with no accuracy, model, validation or false-positive rate given. And he says load and performance testing has become simpler, with tools simulating multiple users hitting an endpoint.

He also says the team trialled several commercial tools for evaluating AI-generated responses and rejected them all, deciding to build its own internal evaluator instead. No vendor is named, and none should be inferred.

The Document-To-Code Pipeline

The first tool is a generative feature-document builder that pulls from Confluence, Jira tickets, test documentation, Figma files and assorted formats including PDFs, slide decks and notes. He frames its value narrowly, as reducing the effort of gathering information about a feature.

The second takes that generated document plus user stories or requirements and produces mind maps in natural language and manual test cases. He narrated a slide with requirements on one side and generated cases on the other, output as a CSV file.

Third comes an internal code generator, which he calls part three of the process. It takes the manual test cases from the previous step and writes the actual test code in a supported language, naming Python and JavaScript, compatible with Cypress, Playwright and Robot Framework.

His claimed advantages for generated tests are that they are faster, accurate and error-free, with no need for the automation skills previously required. He undercuts that himself moments later by noting that usually there are some minor adjustments needed, and he makes similar absolute no-human-error claims about two other tools. Both halves belong in the record.

The design of the chain is the actual argument. Each tool consumes the previous tool’s output, so a documentation page and a design file eventually become executable test code without anyone retyping the intermediate artifacts.

Dashboards, No-Code And Jira

The fourth tool is an engineering insights dashboard consolidating information from Confluence, Jira and a test management tool, with real-time data and links to issues. He says managers value not having to navigate between tools, and not needing a licence for each one.

The fifth is a no-code automation tool driven by a graphical interface and visual workflows, with a library of preconfigured actions, workflows and validation steps that the team built. He speaks generically about the tool itself, so whether the platform is in-house or third-party is not established. The two examples shown were building blocks for clicking an element and entering text into a field.

His stated benefit for no-code is participation rather than speed alone, making it easier for less technical people and QA staff to take part in the testing process. The sentence naming those groups is damaged in the captions and is paraphrased here.

The sixth is an API integrating Robot Framework with Jira. The prerequisite he names is a Jenkins pipeline triggering execution: the job runs, Jenkins generates a log, and the API uploads tests to a Jira test run, creates a cycle, attaches it to a user, and marks each test pass or fail.

The seventh is generative ticket creation, and it carries the clearest worked problem in the talk. Raising a bug normally means converting automation code into human-readable reproduction steps by hand, line by line, so that whoever analyses the issue does not have to read test code. Their tool sends the failing sections of code to a model and asks it to write both the ticket and the reproduction steps.

Counting the AI-response evaluator mentioned earlier, that is eight tools rather than the seven the chapter list numbers. He never numbers them himself.

The Numbers On The Last Slide

End-to-end test design fell from two to three hours down to fifteen to thirty minutes, per his closing metrics slide.

API test generation fell from two to three hours per endpoint to the same fifteen to thirty minutes.

Root cause analysis fell to ten to thirty minutes. He self-corrects mid-sentence here, starting to say debugging before switching to root cause analysis, and the before-state he gives is a bare number followed by two and a half hours, with the unit on the lower bound not audible. The published chapter supplies minutes; the recording does not.

Time to first pull request review came down to five minutes, against a baseline he gives only as a lot of hours.

On money he gives direction and no figures. The team is saving many hours and a lot of cost each year, and he says the total annual model cost is very small. No spend number, headcount, team size, measurement window or date accompanies any item on the slide, so these ranges cannot be turned into percentages or multiples.

His closing thesis restates his opening one: better QA produces better products, better quality and better customer experience. No questions were taken, and the host closed the session immediately afterwards.

This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.

Author

...

TestMu AI

Blogs: 232

  • Twitter
  • Linkedin

TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests