Two records, not one
Disability is the sharpest test case for a claim made throughout this manual: that AI’s effects are neither uniformly good nor bad, and which one shows up depends on who built the system, who deployed it, and what recourse exists when it fails. For disabled people, both records are already real and dated, not hypothetical. Multimodal models have given blind and low-vision users a description tool that did not exist three years ago. The same underlying pattern-matching — inferring things about a person from voice, movement, or typing behavior — is also the subject of a federal discrimination charge, a certified collective-action lawsuit, and a Department of Justice inquiry into a child-welfare algorithm.
This page covers documented, dated cases only. It deliberately excludes one widely circulated 2026 story about Meta account lockouts tied to content-moderation false positives, since the underlying reporting found that to be a general moderation failure, not a disability-specific one. See civil rights, discrimination, and AI surveillance for the legal framework governing the cases below, and labor and economic displacement for automated hiring’s place in the wider employment picture.
What is actually working, and how it’s been validated
The most-cited success is Be My AI, built into the Be My Eyes app on OpenAI’s GPT-4V and launched in March 2023 as the first third-party application built on that model (TechCrunch, March 14, 2023). It lets a blind or low-vision user photograph a scene, a label, or a piece of mail and receive a spoken description. In one deployment, as a support-line supplement inside Microsoft’s operations, Be My Eyes and Microsoft reported a 90% successful-resolution rate and an average resolution time of roughly four minutes, less than half the typical human-agent time (Be My Eyes). That figure is company-reported from one partner deployment, not independently audited — a vendor’s own account of its own product. The same source documents the tool’s most consequential failure mode, confident hallucination: beta testers found Be My AI would describe objects that were not actually present, with the same confidence as accurate descriptions. For a blind user acting on a description — reading a medication label — an undetectable hallucination is a safety problem, not a UX bug.
Voiceitt, which builds speech recognition tuned to atypical and dysarthric speech, ran a pilot with Deaf users that it reported achieved roughly 8% word error rate — better than 90% accuracy — after around 200 training recordings per user, reportedly outperforming mainstream speech recognition for these speakers even before personalization (Hearing Review; PR Newswire). This is a single company-validated pilot, not an independently replicated trial, though a peer-reviewed methods paper separately describes Voiceitt’s training pipeline (Assistive Technology, 2024, DOI 10.1080/10400435.2024.2328082). Google Project Relate takes a similar approach — recognition and synthesis personalized to speech from people with ALS, cerebral palsy, Down syndrome, Parkinson’s, stroke, and traumatic brain injury, needing roughly 500 phrases — and entered closed beta in November 2021 (TechCrunch). No independent efficacy figure for Project Relate could be confirmed; what is public describes the training approach, not a validated outcome.
The pattern is consistent: real products, real use, real numbers — but numbers reported by the companies that built the tools, not independently audited. That does not erase the gains — for a blind user, the comparison is not a validated gold standard but no description at all — but it means treating the evidence as company-reported, and expecting less reliability than launch coverage implies.
Where the same technology becomes an accommodation problem
The employment side shows how quickly an accessibility gap becomes a discrimination claim once AI moves from an assistive role into a gatekeeping one.
A Deaf, Indigenous job applicant, represented by the ACLU, filed charges with the Colorado Civil Rights Division and the EEOC alleging that HireVue’s AI video-interview platform — used by Intuit — penalized her because its automated speech-recognition and scoring system was not equipped to properly assess her speech patterns. She had requested human captioning as an accommodation during the AI-scored interview, was denied it, and was later rejected with feedback to “practice active listening” (HR Dive). This is a charge, not a finding — an allegation working through an administrative process, with no adjudicated outcome reported. It illustrates the mechanism the EEOC and DOJ warned about in their May 2022 guidance on algorithmic hiring, discussed further on civil rights, discrimination, and AI surveillance: a system validated on speech patterns unrepresentative of disabled applicants can screen them out regardless of intent, and an accommodation request can be refused by a process with no human empowered to grant one.
Mobley v. Workday, filed in the Northern District of California on February 20, 2024, is a larger version of the same claim: a class and collective action alleging Workday’s AI applicant-screening system produced disparate impact by race, age, and disability. On May 16, 2025, Judge Rita Lin granted conditional certification to the age-discrimination collective claim under the ADEA, and a named plaintiff — an applicant with asthma and a cancer history — was separately allowed to proceed on an ADA claim tying his rejections to medical-leave history (Holland & Knight; Norton Rose Fulbright). What that ruling is and is not matters: it lets more plaintiffs join and discovery proceed; it is not a finding that Workday’s system actually discriminated. The case remained in discovery through mid-2026, but the order establishes that a federal court found the disability theory concrete enough to survive dismissal — a materially higher bar than a filed complaint clears alone. See job quality, not job counts for how automated screening also degrades outcomes for applicants who pass through.
When the same pattern reaches child welfare
The starkest documented case moves outside hiring entirely. Pennsylvania’s Allegheny County uses a predictive risk-scoring algorithm in its child-welfare system, now under investigation by the U.S. Department of Justice following complaints that it incorrectly flagged disabled parents as neglectful, in some instances contributing to child removal without independent evidence of neglect. That account comes from the Center for Democracy & Technology and American Association of People with Disabilities’ 2025 report below, which describes the DOJ inquiry as ongoing; no confirmed outcome could be verified beyond that, and it should be treated as an open matter, not a closed finding. The mechanism alleged is familiar from the civil-rights literature more broadly: a model trained on historical case records can encode caseworkers’ prior assumptions about disabled parents’ fitness, and a risk score presented as objective can make an old bias look like a neutral number.
The most comprehensive account of this pattern across sectors is CDT and AAPD’s joint report, Building a Disability-Inclusive AI Ecosystem: A Cross-Disability, Cross-Systems Analysis of Best Practices, published March 11, 2025 (CDT; full PDF). It documents harms to disabled people across employment, education, public benefits, and information and communications technology, and is the source for the Allegheny County account above. CDT raised a related, narrower warning years earlier, arguing that hiring interfaces built on mouse movement, keystroke timing, video, or voice-pattern analysis structurally exclude disabled applicants regardless of vendor intent — a report generally dated to around December 2020, though its exact date could not be independently confirmed (CDT, “Algorithm-driven Hiring Tools”).
Reading the two records together
None of this resolves into a single verdict on “AI and disability,” and it should not. The same capability — a model that infers meaning from speech, movement, or behavioral signals a person cannot fully control — produces a genuine accessibility gain when it is optional, transparent about its limits, and controlled by the disabled user, and a genuine discrimination risk when it is mandatory, opaque, and controlled by a gatekeeper who has outsourced judgment to a score. Be My AI, Voiceitt, and Project Relate are tools a disabled person chooses and can stop using. HireVue’s scoring, Workday’s screening, and Allegheny County’s risk model are systems applied to a disabled person by an institution, with consequences they cannot decline.
That distinction — assistive versus adjudicative — predicts harm better than any claim about a model’s sophistication, and it runs through civil rights, discrimination, and AI surveillance: the standard for a rights-affecting system is not whether it is accurate on average, but whether it was validated for the population it screens, whether an applicant can contest an error, and whether a human with real authority sits at the decision point. Every claim of accommodation denial or disparate impact above traces to the same missing element — a validated process, applied to the affected population, with a genuine appeal. Every accessibility gain above traces to the opposite: a tool the affected person controls, with a disclosed failure rate. More capable multimodal systems could extend the Be My AI pattern further. Nothing rules that out. But nothing shows it arriving automatically either: the same capability is exactly what a screening vendor would also acquire.