Hero Background

Next-Gen App & Browser Testing Cloud

Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

Next-Gen App & Browser Testing Cloud
Testμ

Developing AI Agents for Disability and Self Determination [Testμ 2026]

Dr. Lisa Dieker on co-designing AI agents with disabled users, why Project RAISE dropped facial tracking, and why compliance is not the same as accessibility.

Published on:

A research team built an AI agent for children identified with autism, and were pleased with it. The children said it looked stupid.

That, as the researcher tells it, was not on the to-do list. The fix was to stop deciding: around a thousand children in the school voted on roughly six designs, with the scoring weighted toward the students the agent was actually for.

At Testμ Conf 2026, Dr. Lisa Dieker, Williamson Family Distinguished Professor in Special Education at the University of Kansas, walked through Project RAISE and the difference one word makes. She describes herself somewhat differently on air, mentioning 20 years at the University of Central Florida as a Lockheed Martin Eminent Scholar, patent holder, and mother.

Youtube thumbnail

If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.

TL;DR

Co-design is the practice of building a tool with its disabled users from the start rather than making it accessible afterwards. It matters because compliance and accessibility are not the same thing: a tool can meet every standard and still let someone take a survey without being able to create one.

  • What was Project RAISE? - A $2.5 million US Department of Education-funded project, now ended, in which Dr. Lisa Dieker’s team taught children identified with autism to code a robot with help from an AI agent. United Cerebral Palsy of Central Florida led it, and by its final year it covered 10 schools and 150 students in a randomised control study.
  • What does “built with, not for” mean in practice? - Disabled users evaluate and improve a product before it ships rather than receiving an accessible version later. Dr. Lisa Dieker’s argument to builders is commercial as well as ethical: designing for the people with the most complex needs produces a better product for everyone else.
  • Did Project RAISE use facial tracking? - No, not after the first attempt. The team asked its vendors who the models had been trained on and could not learn anything about disability, age or physical difference, so they dropped it and used a simple halo around the face instead.
  • Did the researchers choose what the agent looked like? - No. The first agent the team built was rejected by the children as looking stupid, so they ran a competition in the school where around a thousand children chose from about six designs, weighted toward students on the autism spectrum. The winner became ZB.
  • Where did the agent’s language come from? - From the students. Dr. Lisa Dieker’s team recorded what children actually said while coding a robot, summarised the transcripts in NVivo, and used that as the agent’s script, so it spoke the vocabulary of third graders rather than adult-authored lines.
  • Could the agent tell a child “great job”? - No. The agent had no reliable view of what the child was doing, so performance praise was ruled out. It offered affirmation and executive-function prompts instead, including “you got this”, “stay on task” and the redirection line “back to the iPad”.
  • How did heart-rate data change the agent’s behaviour? - Students wore biometric devices, and when stress rose the agent moved to a tighter cadence and gave more feedback rather than less, including a prompt to take a deep breath. Dr. Lisa Dieker reports the extra feedback did not turn out to be distracting.
  • What were the three phases of the study? - Learning, teaching and integration. Students first learned to code the robot with the agent, then brought a peer who could not code and taught them, putting the student with a disability in the role of expert, and only then did the agent enter the classroom to support self-regulation and communication.
  • What results did Project RAISE report? - Dr. Lisa Dieker reports significant increases in student communication, which she says has been published, plus increases in on-task behaviour and gains in maths practices. She cites no paper, journal or effect size, so these are her account of the findings rather than a citable result.
  • Which disabilities are hardest to design agents for? - Deafness and blindness as single categories, because agents lack good enough lips for lip reading and lack digits to sign, while a blind user may not care about an agent’s appearance at all. The harder case is multiple disabilities, where teams overdose a person with an all-purpose agent instead of finding the specific gap.
  • Should autonomous AI agents be used with young children? - Not below third grade, which Dr. Lisa Dieker presents as her own rule rather than a research finding. The Project RAISE agent was closed-loop, unable to learn or act on its own, and on genuinely autonomous agents her answer was that she does not know and that is when she gets scared.
  • Does meeting accessibility compliance mean a tool is accessible? - No, and this is her sharpest point. She cites tools that meet compliance standards yet let you take a survey without creating one, read only the X and Y axes of a chart without the data, and sometimes will not read aloud at all.

Built With, Not For

The host opens with the framing the whole talk hangs on. Most technology built for accessibility is built for disabled people and far less of it is built with them, and that single word is the difference between a tool that technically complies and one that works in real life.

Dieker takes the framing and repeats it in her own terms. Technology for disability is too often retrofitted rather than designed with disabled people in mind from the start.

Her stated motive for a special education professor being deep in technology is that technology can give equity to people with disabilities, and it is a tool that makes a measurable difference.

Her aim for the audience is practical rather than moral. Co-designing with disabled users will save you time and money, because the group you design for last is the group that forces the expensive rebuild.

She reframes disability design as majority design. Teams design for the 90%, the people she works with are in the 1%, and designing for that 1% produces a better product for the other 90%, because in her words access often stinks.

Josh And Self-Determination

Her first evidence is personal. Her son Josh, now 30 and recently married, was born with Tourette syndrome, dyslexia and dysgraphia, and was at the same time the 20th best gymnast in the United States. Strengths and deficits in the same person.

Her formulation of the lesson is that it was never a question of whether he could succeed, but whether the people around him could provide the supports he needed for success. She believes AI agents are the next form of that support.

The rule she draws is a question for builders rather than a principle. Self-determination begins with voice: have you asked disabled people what they think of your tool, have you put them in focus groups, and are you really listening?

She adds a commercial argument to the ethical one. Whatever disabled users tell you about your tool is what the average older or less tech-savvy user will want to know too, so the feedback generalises well beyond the group you collected it from.

The supports that mattered for Josh were concrete and mostly pre-AI, including Bookshare, the US service providing digital books free to people with print disabilities, which she calls a game changer. She names Copilot and NotebookLM as the current generation and expects agents to be the next gap-fillers, as personalised coach, self-advocacy support and transition support.

AI As Cognitive Prosthetic

She rejects the everything-agent. Not everybody needs an agent for everything, and there are things she is better at doing herself, though she calls AI her second brain for the days she is tired and does not know what to cook.

For disability specifically she uses a term she credits to leaders in her field, naming Jamie Basham among them: the cognitive prosthetic.

The prosthetic framing carries a limit she is explicit about. Agents should fill gaps in things a student is not good at, and must not take away the core concept of learning, the same way a wheelchair does not remove the person.

She grounds this in supervision experience rather than theory, having graduated 24 PhD students with disabilities spanning blindness, deafness, cerebral palsy, ADHD, autism spectrum and traumatic brain injury.

What those students taught her is the section’s punchline. The things being built in AI are not really very accessible even though the builders think they are, because teams meet minimum standards and not the standard for someone with multiple disabilities.

She flags invisibility as a design problem too. Not every disability is visible, and if you met her son you would not know he has Tourette syndrome, yet it plays a part in his brain every day.

The Accessibility Paradox

Her paradox: AI makes the world more accessible, translating, summarising, simplifying text and handling speech, while the interfaces wrapped around it get less accessible. She uses the session itself as the example, noting that if she were blind, the deck she is presenting raises questions about what is on screen that she could not answer alone.

Prompting is her second example. The literacy demand of writing good prompts is heavy and inaccessible to people who read below a fifth-grade level. Outputs fail the same way, and she asserts that beautifying slides in a mainstream tool instantly makes them inaccessible to a screen reader.

Her stated greatest pet peeve is representational rather than technical. Ask an AI system about disability and it renders somebody in a wheelchair every time, erasing the actual diversity of disability. Her summary line is that use does not equal design: disabled people use these tools and rarely get to design them.

She puts a number on the gap, saying disability representation in AI design is 16% globally. The figure arrives on a slide with no source, year or definition of what is being measured, so it is her framing rather than a citable statistic.

Her stake in it is concrete. Disabled people are already often unemployed or underemployed, so if AI disrupts the workforce, that group absorbs the disruption first.

The two technical gaps she names are reliable multimodal access, meaning captions that miss names including her own surname, poor image descriptions, and readers that announce the axes of a chart without ever conveying the percentages; and cognitive access, which she defines as the ability to translate content down rather than only across languages. Build an agent that writes at a 14th-grade level and translation becomes a reading-level problem, which she does not think is in builders’ minds.

Retrofit Design Assumptions

She walks four everyday objects to show what a wrong design assumption costs. The microwave was revolutionary unless you are blind, with flat buttons and no braille, so now an app exists to tell a blind user what is on the panel, when raised buttons at the design stage would have made the app unnecessary.

She extends that case to ageing rather than disability alone, citing a figure for sight loss after 80 that she gives without a source, and asks why accessibility is an after-the-fact patch rather than a starting assumption.

Cars assumed a driver, and she points to autonomous vehicles and lidar as the rethink of mobility, though her point is that the original design assumption had to be undone.

Phones assumed small screens, fixed buttons and a sighted user, and yet, by her own concession, you cannot now find a better device for accessibility features than a handheld. Those features were added after the fact rather than designed in.

Computers still require adding gadgets to make participation possible, and she estimates, explicitly as an estimate, that around 20% of users would benefit from bigger mice, better buttons and better colour schemes that are rarely considered.

The bridge to her research is the shift from build-it-then-fix-it to co-design from the start, because disabled people should not have to wait for the accessible version. In the flipped model, disabled users evaluate and improve a product before it goes to market, which she argues also means more customers from day one.

Note

Note: Accessibility that is designed in beats accessibility that is patched on. Try TestMu AI now!

TeachLivE Avatars

Before Project RAISE, Dieker worked on TeachLivE, which she describes as the first simulator to train teachers with avatars, a class of middle-school avatars that can act up, laugh or act confused while a teacher practises. She holds the patent with colleagues and it is now licensed commercially, so the claim to being first comes from an interested party.

When the team wanted to add students with disabilities to the simulator, the lesson was that those avatars had to be co-designed with real people rather than authored.

Martin, one of the avatars, is modelled on a real young adult student with autism, looks exactly like him, and his movements and actions were built with him and his family.

She states the limit of that likeness in the field’s own phrase: if you meet Martin, you have met one person with autism, because he is one person with autism.

Bailey, the avatar beside him, was created with a young woman with Down syndrome in the family, and the family asked that Bailey not look like her, so teachers could practise working with students with intellectual disabilities without importing assumptions from her appearance.

Dieker admits she would have chosen differently, which is the co-design principle applied to herself: she might have taken another path, and these families helped the team understand what mattered from their viewpoint.

Inside Project RAISE

Project RAISE was a US Department of Education-funded $2.5 million project, now ended, whose job was to teach children specifically identified with autism to code a robot, and to build an AI agent that would help them do it, well before building an AI agent was fashionable.

It started right as COVID hit, with United Cerebral Palsy of Central Florida as lead partner and Dieker’s team as research lead, working in a school serving children with multiple disabilities.

The first mistake is the one she leads with. The team built an agent they were excited about, and the children said it looked stupid, which she notes was not on the to-do list.

The fix was to hand the decision over. The team ran a competition in the school where roughly a thousand children picked from about six agent designs, with scoring weighted toward students on the autism spectrum because they were the target group. Her own counts here are hedged rather than exact.

Language was sourced the same way rather than written. The team recorded what students actually said while coding a robot, ran the transcripts through NVivo to summarise them, and fed that back so the agent spoke the authentic language of third graders who struggled with disabilities.

The mechanism that made this possible was human-in-the-loop puppeteering borrowed from the earlier avatar work. A human drove the agent first, which is how the team learned what children valued before automating any of it.

Dropping Facial Tracking

Project RAISE began with facial tracking for a narrow purpose: detecting when a child had disengaged and was no longer looking at the agent.

The team stopped when it could no longer feel confident the tracking would not be biased against the very students it was being used on, after vendors could not describe the disability, age or physical composition of their training data.

The replacement was deliberately cruder. A halo around the face rather than facial tracking, enough to know whether the child was present without inferring anything from their expression.

The team also abandoned emotion inference. Early designs gave feedback based on whether the child appeared happy or sad, and in her words it was not only messy, they did not feel it was ethical at the time.

What the agent said instead was behavioural and neutral. “Back to the iPad” was a redirection when a student went off task, making no claim about the student’s internal state.

She generalises the decision into a question for the audience: are you sure you are creating the right data set, and are the people it describes part of the decision-making?

Next-generation test execution with TestMu AI

ZB, RAISY, Three Phases

The agent that survived the competition is ZB. RAISY is the virtual robot the students learn to code with Blockly, and ZB’s job was to help them code RAISY. Dieker stresses that ZB’s job was very narrow and completely directed at these students’ needs.

The stated goals were basic coding, increased communication skills and increased executive functioning. The finding she highlights is that the narrow agent turned out to be good for everybody in the classroom as it was built out over five years.

PhaseWhat happens
One: learningResearchers observe students with their real teacher, then students learn to code RAISY around a basic square with ZB as their friend
Two: teachingThe student with a disability brings a peer who cannot code and teaches them, at a standalone station with no teacher required
Three: integrationZB comes into the classroom to help students self-regulate and communicate, observed over six weeks

Phase two inverts the usual role, putting the student in the position of being a giver instead of the one who always needs help. Dieker notes teachers loved that part.

The cycle repeated across five years as a design grant, where each year’s outcomes had to improve and new schools had to be added. By the final year the study had grown to 10 schools and 150 children in a randomised control study, with almost a thousand students passing through over the life of the project.

Biometrics And Affirmation

The team connected ZB to biometric devices worn by students so the agent could see a child’s heart rate, and Dieker is candid that they did not know how this would go.

The agent ran on a cadence rather than reacting to events. When a student’s stress rose, ZB moved to a tighter cadence and gave more feedback rather than less, including prompts such as taking a deep breath.

Crucially the agent could not praise performance, because it had no reliable view of what the child was doing at that moment. “Great job” was ruled out.

What replaced praise was affirmation and executive-function prompting drawn from the human-in-the-loop phase: you got this, stay on task, I believe in you.

Dieker reports that the added feedback was not distracting, and that the study found significant increases in student communication, which she says has been published, along with increases in on-task behaviour and gains in maths practices. No paper, journal or effect size is cited in the session.

She flags her own uneasy finding rather than hiding it. Students would say they were being good for ZB, and the team had to reckon with the fact that he is not real. She attributes the attachment to the agent speaking the students’ own recorded language.

Support Decisions, Don’t Make Them

Her forward list for agents is goal setting, self-advocacy, job interviewing, transition planning and community life, with the constraint that agents support people to make their own decisions rather than deciding for them. Self-determination is the core value of her field.

Addressed to builders: if she were building agents for people like her son, she would build reading support, movement support, self-advocacy and college-admission support that does not hand over the answers. “I hate chemistry” should be met with an offer to practise chemistry, not with the chemistry answer.

Her framing of the design error is that teams build the agent they think humans need instead of agents that align with the human in front of them.

On employment she flips the accommodation model. Instead of hiring someone and then retrofitting the curriculum or the workplace, have the agents already available when they arrive, as part of the DNA of the business.

Her current project, running through September 2027, applies the same co-design rule to teachers. Teachers with disabilities are on the design team, and one of the project’s coaches is blind and has shaped how the coaching cycle and observation tool work, including how she uses and tags audio.

The observation tool layers multimodal data, with a coach’s tags timestamped against the teacher’s biometric device and facial analysis. The design questions she says improve the product are about edge cases: what if someone has an irregular heart rate, or cerebral palsy affecting their facial movement.

Compliance Is Not Accessibility

The newest artefact she mentions is FLITE ICA, a patent she says was granted the Monday before the session. It captures audio, analyses it and timestamps everything in under two minutes so a coach can point to a moment and say more pause was needed, surfacing language patterns, voice modulation and facial cues. That is a days-old claim about an unreleased research tool with no independent validation shown.

Where she wants to take it next is the neurophysiological state of teacher mind and body, then the same for students, on the reasoning that if you know what a mind and body are missing in the moment, an agent can prompt someone to stay on track or breathe.

Her research dream is agents that detect overload, support self-regulation, adapt as stress increases and help a person advocate sooner rather than after the fact, conditional on disabled people being in the design from the start, biometric and facial data especially.

Her hardest line is aimed at standards. Tools that have met compliance standards still let you take a survey but not create one, still read only the axes of a chart, and sometimes will not read to you at all. They have met compliance standards, and it is still a huge issue.

On academia she is blunt from experience, saying that if you want to get a PhD and you have a disability, good luck, because the world is still awful. That is her view rather than a finding.

Comma

Q & A Session

Seven audience questions closed the session. Two of them she did not answer as asked, and both are worth reading as she left them.

  • If an AI agent acts like an employee, should we secure it like an employee, like an application, or as something new?

    Dr. Lisa Dieker: If an agent is meant to act like an employee and makes mistakes, you should have designed a way to give it feedback, the way you would an employee. Another speaker at a conference said he made 150 agents and fired 74, and the instinct is right: if a tool is not working for you, get rid of it. Of the 150 children in our final year, all liked ZB except one, who fired the agent in the first week, and that was the student’s right. Her answer is about management rather than security, and never addresses security controls.

  • How do you design agents that support independence without creating dependency?

    Dr. Lisa Dieker: Not everybody uses the same wheelchair, so it starts with personalisation rather than with limiting use. If I could walk and stayed in a wheelchair because I liked it, that would still be my choice. And everybody is trying to sell me an agent that is not the one I need. She resists the premise that dependency is the designer’s call to make, and her sharpest line lands on the market rather than the user.

  • Which kinds of disability have been most challenging to develop for?

    Dr. Lisa Dieker: Two populations. Agents are poor for deaf users, because they do not have good enough lips for lip reading and do not have digits to sign, and the field is not building with that in mind. A blind user can use an agent but may not care about its physical features at all, which is the opposite risk, overdesign, and we spent a great deal of time on how ZB looked. Harder still is multiple disabilities, where you overdose a person with an all-purpose agent instead of finding the specific gap.

  • If there is a security incident, who is responsible behind an autonomous AI agent?

    Dr. Lisa Dieker: My rule is that below third grade you should not go there, which is why Project RAISE targeted third grade and above with a closed-loop system that could not learn and could only do what we allowed. On autonomy I do not know, and that is when I get scared. Schools and mental health are the two domains requiring the most care. She does not answer the accountability question and effectively says so, and no liability model is offered.

  • What ethical lines come up, and how do you handle them?

    Dr. Lisa Dieker: My first ethical line is representation: not enough women and not enough disabled people in the field, so the technology is created by a population that does not represent the population it serves. My second is guardrails nobody understands, not the parent, not the teacher, sometimes not even the creator. This feels like the early internet. She redirects away from privacy and consent entirely, and contrasts agents unfavourably with the guardrails she attributes to autonomous vehicle companies, an assertion she offers without support.

  • Could you use the word “challenged” rather than “disability”?

    Dr. Lisa Dieker: I love the sentiment. But in the United States, disability is the word by which services are funded, supported and identified, and it is the federal guideline. The term I prefer is unique abilities. And a challenge can pass, whereas a disability never goes away and is part of your life forever. I am not in charge of the federal law, or it would already have changed. She declines the substitution on practical grounds rather than on principle.

  • What would you recommend for designing agents for people with varying disabilities, such as senior citizens?

    Dr. Lisa Dieker: There is going to be a large population of older people who will not be able to use what is being built, for two reasons: they are not being trained fast enough, and the products are not being made easy enough to use. Build to a standard. If everything built met WCAG 2.2, and those standards were expanded to cover elderly users, we would reach a new level. Design across the full range of disability and you are accessible for the hundred-year-old as well as the two-year-old.

This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.

Author

...

TestMu AI

Blogs: 232

  • Twitter
  • Linkedin

TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests