Lab Index
Every example in the PAN Lab
127 organizational networks you can explore, stress test, and govern in PAN Lab v0.1. Each has its own page: what it models, the sources and evidence behind it, and the concepts, pressures and levers it connects to across the Governance Center.
Every context is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Some of these organizations depend on the same outside supplier, reviewer or assessment instrument as another one in this list. Common Cause reads across the index for those shared dependencies, and quotes the evidence for each one.
Every pressure the Lab can apply has a page of its own, with what it pushes on and which levers answer it. Read about the Lab's pressures
Every lever the Lab offers has a page of its own, with what it pushes on and which pressures it answers. Read about the Lab's levers
Compare approaches
Teaching networks that hold the model fixed and vary only governance — not models of one real deployment.
The same AI running hands off: the agentic office
In this made-up example office, AI agents act on cases themselves, and a stretched staff waves most of their work through.
1 cited claim · 12 levers · Details · Open in the Lab
The same AI with a human checking: the supervised office
In this made-up office, staff review an AI assistant's drafts before anything enters the case records. The risk is that review slowly becomes a formality.
1 cited claim · 12 levers · Details · Open in the Lab
The same AI under full guardrails: the professional office
A modeled office, not a real one: an AI assistant under strong checks, tested by staff pasting client data out and a silent vendor update.
1 cited claim · 12 levers · Details · Open in the Lab
Behavioral-health & crisis triage
10 examples
REACH VET
The flag that works and the outcome it misses: a suicide-risk model
Each month REACH VET, a Veterans Affairs model, flags high-risk veterans for suicide-prevention outreach. Appointments rose, but two evaluations found no drop in suicide deaths.
2 cited claims · 12 levers · Details · Open in the Lab
Vanderbilt VSAIL suicide-risk alert
The alert that had to be dismissed: an EHR suicide-risk model
In a trial, Vanderbilt's suicide-risk alert led clinicians to screen in 42 percent of flagged visits as a pop-up, 4 percent as a chart icon.
2 cited claims · 11 levers · Details · Open in the Lab
Kaiser Permanente Suicide-Risk Model
The added sensor: an EHR-embedded suicide-risk score
Kaiser Permanente's machine-learning score flags intake patients at risk of a suicide attempt. The flag, or a positive questionnaire, prompts a clinician to assess them.
2 cited claims · 9 levers · Details · Open in the Lab
Crisis Text Line & Loris.ai
The corpus and the spinoff: governing crisis-conversation data
Crisis Text Line, a crisis texting service, ranks waiting texters with a model. Loris.ai, a for-profit it partly owned, trained software on their conversations.
2 cited claims · 10 levers · Details · Open in the Lab
NarxCare
The score you cannot see: an opaque prescribing-risk model
Bamboo Health's NarxCare turns state prescription records into patient risk scores. Patients cannot see or contest the scores, which can gate access to pain medication.
2 cited claims · 13 levers · Details · Open in the Lab
Limbic Access (NHS Talking Therapies)
The front door that works: a self-referral triage chatbot
The Limbic Access chatbot takes NHS talking-therapy self-referrals. Referrals rose more in services using it. Nearly all evidence it works comes from its maker.
2 cited claims · 9 levers · Details · Open in the Lab
Woebot (a governed app wind-down)
The responsible wind-down: retiring a peer-reviewed CBT chatbot
In 2025 Woebot Health retired Woebot, its therapy chatbot app, on a published schedule. Users could request their chat transcripts before account data was anonymized.
2 cited claims · 9 levers · Details · Open in the Lab
LyssnCrisis counselor QA at ProtoCall Services (988)
The AI watches the counselor rather than the caller: a crisis-line QA scorer
LyssnCrisis rates how counselors on 988, the US crisis line, handle calls, never the callers. No outside study has checked its scores or their benefit.
2 cited claims · 10 levers · Details · Open in the Lab
Gaggle Safety Management
The discontinuation that changed the vendor rather than the graph: a school monitoring scanner
Gaggle's software scans students' school accounts for signs of self-harm. When students sued, one district switched vendors. Did the practice change, or only the name?
2 cited claims · 11 levers · Details · Open in the Lab
Oxevision camera monitoring on NHS mental health wards
The evaluation was written by the seller: a bedroom monitor no one independent checked
NHS mental health wards film patients' bedrooms with Oxevision cameras. One hospital trust's procedure set consent aside, and the seller shaped the evidence for it.
2 cited claims · 10 levers · Details · Open in the Lab
Benefits navigation & public-facing chat
11 examples
Nava assistive benefits chatbot
Done carefully: a verify-before-use copilot
Nava's chatbot answers caseworkers' benefits questions from vetted documents, with quotes to check. The question is what keeps this cautious design in place under pressure.
10 levers · Details · Open in the Lab
GOV.UK Chat
The gate that said not yet: a public assistant behind a staged pilot gate
GOV.UK Chat answers the public's questions from official guidance. Its gate held back a version that fell short, but its builders grade its accuracy themselves.
2 cited claims · 10 levers · Details · Open in the Lab
Burokratt
The network of networks: a federated public-service chatbot
Estonia's Burokratt puts many public institutions' chatbots behind one window. A mistake in its shared routing or knowledge can appear in every institution's answers.
1 cited claim · 11 levers · Details · Open in the Lab
Caddy adviser copilot at Citizens Advice
The gate is a job rather than a habit: a supervisor-checked adviser copilot
Caddy drafts benefits answers for Citizens Advice advisers, and a supervisor checks every draft before an adviser sees it. Can that check survive wider use?
2 cited claims · 10 levers · Details · Open in the Lab
Albert France Services
Killed without a number: a sovereign adviser assistant that no metric ever measured
Albert, a French state AI assistant, drafted answers for France Services advisers in a pilot. No error rate was published, and it was not extended.
2 cited claims · 10 levers · Details · Open in the Lab
Benefits Data Trust wind-down
The node that could not be kept: winding down a benefits-navigation nonprofit
In June 2024 Benefits Data Trust's board voted to close the benefits-navigation nonprofit within 60 days. Its government partners had no designated successor.
2 cited claims · 10 levers · Details · Open in the Lab
Propel in-app SNAP benefits assistant
Read-only by design: a benefits assistant grounded on the record it never writes
Propel's app uses AI to help people learn why a SNAP food-benefit deposit did not arrive. It reads the state's record, never writing to it.
2 cited claims · 9 levers · Details · Open in the Lab
GetCalFresh
Most of a state's online intake with no authority at all: an assisted-application node and its handoff
GetCalFresh, a nonprofit's online form, carried most of California's online food-benefit applications but decided none. What if a service no law required carries that much?
2 cited claims · 10 levers · Details · Open in the Lab
Frida (NAV Norway)
Ask for a human: the handover boundary as a governed surface
Frida is the chatbot of NAV, Norway's welfare agency. The case asks how easily citizens reach a person and who reads chats ending without one.
2 cited claims · 11 levers · Details · Open in the Lab
IRS collection chatbots
Expanded without a ruler: a federal collection chatbot with no performance measures
The Internal Revenue Service (IRS) expanded its collection chatbot and live chat, and made live chat permanent, with no performance measures or reliable statistics.
2 cited claims · 10 levers · Details · Open in the Lab
Singapore's chatbot fleet refresh
Eighty engines into one: a whole-of-government chatbot fleet refresh
Singapore set out to move its agencies' chatbots off scripted engines, which match keywords to prewritten answers, onto a few shared AI engines.
2 cited claims · 10 levers · Details · Open in the Lab
Caseworker documentation & copilots
8 examples
Magic Notes (Beam)
The drafted record: a case-notes copilot
UK council social care staff draft case notes, the client's record, with Magic Notes, AI software from the company Beam. They review each draft first.
10 levers · Details · Open in the Lab
Guided walkthrough: agentic low-oversight office
A walkthrough-only version of the agentic low-oversight office: an AI assistant that also acts as an autonomous agent, a stretched staff supervising many automated actions at once, and — added for the tour — the Outside…
1 cited claim · 12 levers · Details · Open in the Lab
Minute / Local Transcribe
The governed thing and the measured thing: a state-built meeting scribe
The UK government built Minute, an AI meeting scribe, and piloted it in local councils. Its approval checks covered data handling and process, not accuracy.
1 cited claim · 10 levers · Details · Open in the Lab
DWP Whitemail Insights and Vulnerability Scanner
The letter no one reads twice: an upstream vulnerability scanner
The UK welfare department's AI scans about 25,000 letters daily, flagging signs of risk like self-harm for caseworkers. No one rechecks the rest for risk.
2 cited claims · 10 levers · Details · Open in the Lab
Justice Transcribe
The note that scores you: a copilot at the head of a risk pipeline
Justice Transcribe drafts summaries of probation meetings for the case record in England and Wales. The ministry's reoffending-risk tools draw on probation records too.
1 cited claim · 12 levers · Details · Open in the Lab
VA claims automation (automated survivor-benefit decisions)
The record no one read: automated survivor-benefit decisions
The VA's automation grants survivor benefits with no human involvement when its rules match. A 2026 audit found deficiencies in nearly all such grants.
2 cited claims · 11 levers · Details · Open in the Lab
Amsterdam Smart Check
The governed exit: a fair welfare screener that shipped every safeguard but one
Slimme Check flagged welfare applications for investigation. Its bias, reduced on past data, shifted to other groups in a live pilot. Amsterdam halted it.
1 cited claim · 12 levers · Details · Open in the Lab
Massachusetts DTA call summaries
The record and its source: AI summaries of benefits calls
Massachusetts piloted AI summaries of food-benefit eligibility calls. Each summary is saved into the case record, and the full call transcript is not kept.
5 cited claims · 8 levers · Details · Open in the Lab
Child welfare & family services
15 examples
Allegheny Family Screening Tool
The score and the screener: a child-welfare risk tool
The Allegheny Family Screening Tool scores reports to the county's child protection hotline. Hotline staff weigh the score when deciding which families to investigate.
2 cited claims · 10 levers · Details · Open in the Lab
Illinois Rapid Safety Feedback
The alarm that cried wolf: a child-safety scorer
Illinois's child-welfare agency scored children in reports of suspected abuse or neglect. Its tool flagged thousands as extreme risks and missed children who died.
2 cited claims · 13 levers · Details · Open in the Lab
Allegheny Hello Baby
The help that keeps a file: a birth-risk prevention model
Allegheny County's Hello Baby model scores newborns from county records. Highest-scored families are offered voluntary help, not investigation, by workers who must report suspected abuse.
1 cited claim · 10 levers · Details · Open in the Lab
Douglas County Decision Aide
The score read only at the edges: a child-welfare screening aide
Douglas County, Colorado's Decision Aide scores reports of possible child maltreatment (referrals). An independent trial found it sped decisions without significantly changing children's outcomes.
1 cited claim · 10 levers · Details · Open in the Lab
Eckerd Rapid Safety Feedback
Endorsed then evaluated: a risk tool that spread on a claim
Eckerd, a child-welfare nonprofit, built Rapid Safety Feedback to flag cases resembling past child deaths. It spread to several states before anyone independent tested it.
2 cited claims · 11 levers · Details · Open in the Lab
ProKid (Netherlands)
The colour and the record: a child risk-profiler
ProKid, a Dutch police instrument, sorted children under 12 into colour risk bands from police records, including children recorded only as victims or witnesses.
1 cited claim · 9 levers · Details · Open in the Lab
Insight Bristol / Think Family Database
The database nobody could audit: a shared child-risk profiling system
Bristol's council and police scored children from a shared database. Staff distrusted the scores. Two models were withdrawn, reportedly with no record of why.
2 cited claims · 12 levers · Details · Open in the Lab
Sistema Alerta Niñez (Chile)
Scored before anyone knocks: a child-risk targeting tool
Chile's Sistema Alerta Niñez ranks children by risk, from data families gave for benefits. Local offices use it to decide who gets early help first.
1 cited claim · 11 levers · Details · Open in the Lab
Los Angeles County Project AURA
Caught at the gate: a child-abuse risk model that never shipped
Los Angeles County tested Project AURA, a child-fatality risk score, on past cases. It falsely flagged 3,829 children, correctly flagged 171, and was never used.
1 cited claim · 12 levers · Details · Open in the Lab
What Works for Children's Social Care ML pilots
The bar it never cleared: a child-welfare prediction pilot
A research centre's 32 models predicted which children's cases would escalate. None was right often enough to clear its published bar, so none went live.
1 cited claim · 11 levers · Details · Open in the Lab
New Zealand MSD Predictive Risk Modelling
Halted before it ran: a national child-risk model
New Zealand's Ministry of Social Development commissioned a model to score newborns' maltreatment risk. A minister halted its study, and the tool never went live.
1 cited claim · 11 levers · Details · Open in the Lab
Gladsaxe model
The screen that never ran: whole-population child scoring
Gladsaxe, a Danish municipality, built a model to score every young child for vulnerability. It was stopped in development and never ran on live cases.
2 cited claims · 10 levers · Details · Open in the Lab
Hackney / Xantura Early Help Profiling
The pilot that quietly failed: a small council's family profiler
Hackney Council paid Xantura to flag at-risk families, without telling them. Poor data meant the pilot found few new families, so the council dropped it.
3 cited claims · 11 levers · Details · Open in the Lab
Oregon Safety at Screening
The fix then the off switch: a fairness-corrected screening tool
Oregon's child-welfare hotline screeners saw a fairness-corrected risk score before deciding whether to investigate a report. The agency stopped using the tool in 2022.
1 cited claim · 12 levers · Details · Open in the Lab
Predict-Align-Prevent
The map, not the score: a place-based risk surface and the records it concentrates
A nonprofit's model ranks small map cells by child maltreatment risk, scoring no family. In New Hampshire, teams aimed outreach and funding where it pointed.
3 cited claims · 13 levers · Details · Open in the Lab
Clinical decision support & deterioration alerting
6 examples
TREWS sepsis early-warning system
The alert that works only when confirmed: a sepsis early-warning model
TREWS, built at Johns Hopkins, scores inpatients at five Hopkins hospitals for sepsis. Its benefit came only from alerts a provider confirmed within three hours.
2 cited claims · 8 levers · Details · Open in the Lab
Advance Alert Monitor (AAM) deterioration model
The alert that never reaches the bedside: a screened deterioration model
Kaiser Permanente's Advance Alert Monitor flags hospital patients predicted to worsen in about twelve hours. Regional nurses screen every alert before the bedside team acts.
1 cited claim · 8 levers · Details · Open in the Lab
Sepsis Watch deep-learning detection system
The nurse who gets the alert can't give the order: an authority-split detector
Duke University Hospital's Sepsis Watch alerts a rapid response nurse, who cannot order treatment. The alert becomes care only if the nurse persuades a physician.
2 cited claims · 8 levers · Details · Open in the Lab
Epic Sepsis Model
Switched on before anyone checked: a proprietary sepsis model at scale
Hundreds of hospitals switched on the Epic Sepsis Model before anyone independent checked it. When researchers did, it caught about a third of sepsis cases.
2 cited claims · 8 levers · Details · Open in the Lab
nH Predict Utilization Review
The order of operations inside a coverage determination
UnitedHealthcare and naviHealth, both owned by UnitedHealth Group, decide whether members leaving the hospital get skilled nursing or rehabilitation care covered, and for how long.
4 cited claims · 14 levers · Details · Open in the Lab
CA-CDS Child Abuse Alerting
The mandated report: alerting that writes into another organisation
A consortium at UPMC Children's Hospital of Pittsburgh built child-abuse alerts for emergency departments. Clinicians who suspect abuse must report it to a state agency.
4 cited claims · 9 levers · Details · Open in the Lab
Clinical documentation copilots (ambient scribes)
3 examples
Kaiser Permanente ambient AI scribe
The draft becomes the record: a well-governed ambient scribe
Kaiser Permanente physicians used an AI scribe that drafts visit notes. The physician edits and signs each draft, which then enters the medical record.
2 cited claims · 11 levers · Details · Open in the Lab
Ambient scribe RCT + monitoring playbook
The trial and the playbook: an evidenced and monitored ambient scribe
UW Health tested Abridge, an AI that drafts clinical notes, in a randomized trial, and published a playbook for monitoring it in use.
2 cited claims · 11 levers · Details · Open in the Lab
An ambient AI scribe at a multi-specialty health system
One number and two outcomes: an ambient scribe that split by clinician
Sutter Health's AI scribe served two clinician groups. Of primary-care physicians, 85.8 percent reported improved satisfaction, against 36.4 percent of specialists. One average hides that.
2 cited claims · 11 levers · Details · Open in the Lab
Content moderation & editorial AI
6 examples
YouTube Covid-19 enforcement
The reviewers went home and the error rate showed
With reviewers home in the pandemic, YouTube chose to over-remove by machine. Removals more than doubled. About half of appeals succeeded, up from a quarter.
2 cited claims · 8 levers · Details · Open in the Lab
Meta content enforcement
Ninety percent overturned on the cases chosen to be seen
Meta's software makes millions of content decisions. The Oversight Board picks a few, some likely to be wrong, and in 2023 overturned about 90 percent.
2 cited claims · 8 levers · Details · Open in the Lab
Sports Illustrated AI bylines
The byline nobody was behind
Sports Illustrated published product reviews by invented writers, produced by its contractor AdVon Commerce. An outside investigation found the fabrication, not the outlet's own editors.
2 cited claims · 7 levers · Details · Open in the Lab
CNET AI-drafted articles
Half the articles corrected under a byline that promised a review
CNET published 77 AI-drafted finance explainers under a staff byline. Its own audit found that 41 needed correction. The byline implied a human review.
2 cited claims · 8 levers · Details · Open in the Lab
StopNCII & Take It Down
Two ways to run a hash bank: the intimate-image removal pair
StopNCII and Take It Down fingerprint a person's intimate images on their own device, so partner platforms can find copies shared without consent.
5 cited claims · 12 levers · Details · Open in the Lab
YouTube's Content ID copyright matching system
The party that gains answers the objection
YouTube's Content ID checks every upload against rights-holders' reference files. On a match, the rights-holder's preset instruction applies, and that rights-holder answers any dispute.
8 cited claims · 10 levers · Details · Open in the Lab
Customer service & contact-centre AI
4 examples
A contact centre's generative-AI agent assist
Fifteen percent on average and thirty for the newcomers
A Fortune 500 software firm gave support agents an AI copilot. On average they resolved 15 percent more issues per hour. Newer agents gained most.
2 cited claims · 7 levers · Details · Open in the Lab
DPD customer-support chatbot
The guardrails that stopped holding after an update
After a system update, a customer prompted DPD's support chatbot to swear and write a poem calling DPD the worst delivery firm in the world.
2 cited claims · 7 levers · Details · Open in the Lab
Air Canada chatbot
A policy the chatbot invented and the company that answered for it
Air Canada's website chatbot told a customer they could claim the reduced bereavement fare after booking. The airline refused, and a tribunal held it liable.
2 cited claims · 7 levers · Details · Open in the Lab
Klarna AI assistant
Two-thirds of chats handled and a year later a rethink
Klarna said its AI assistant handled about two-thirds of customer-service chats in the assistant's first month. About a year later, Klarna said quality had fallen.
2 cited claims · 8 levers · Details · Open in the Lab
Education AI
3 examples
Chicago Public Schools On-Track indicator
The rule a teacher can explain
Chicago Public Schools uses a readable rule, not AI, to flag ninth graders as off track to graduate. Graduation rates later rose to record highs.
2 cited claims · 7 levers · Details · Open in the Lab
Wisconsin DEWS
Wrong most of the time and unevenly by race
Wisconsin's Dropout Early Warning System labeled students 'high risk'. An audit found most graduated anyway, and Black and Hispanic students were wrongly labeled more often.
2 cited claims · 10 levers · Details · Open in the Lab
Cleveland State remote proctoring
A camera in the bedroom and a flag that lands by skin tone
Cleveland State University required a webcam scan of students' rooms before online exams, and proctoring software flagged suspected cheating. A court held the scan unconstitutional.
2 cited claims · 8 levers · Details · Open in the Lab
Hiring & employment screening AI
4 examples
Amazon recruiting engine
It learned who was hired rather than who succeeds
Amazon's experimental resume scorer, trained on ten years of resumes mostly from men, penalized "women's". Amazon scrapped it when edits could not guarantee a fix.
2 cited claims · 8 levers · Details · Open in the Lab
Workday AI screening
One model across thousands of employers and testing no one can see
Workday's AI screens applicants for thousands of employers. A lawsuit alleges age discrimination, and a court ruled Workday need not hand over its bias testing.
2 cited claims · 8 levers · Details · Open in the Lab
Unilever and HireVue graduate hiring
Good audits with real savings — and a cohort no one can see
Unilever screened recent graduates with two vendors' AI tools: pymetrics games, then HireVue video-interview scoring. Its reported gains cover only the applicants it advanced.
2 cited claims · 7 levers · Details · Open in the Lab
Intuit's recorded video assessment for promotion
Assessing an incumbent: a recorded promotion gate and the captioning request
Intuit put a Deaf employee's promotion through a recorded HireVue video assessment. A pending charge alleges Intuit denied the human captioning she asked for.
7 cited claims · 12 levers · Details · Open in the Lab
Housing & homelessness services
11 examples
Allegheny Housing Assessment
The score and the scarce bed: a coordinated-entry housing tool
Allegheny County scores people experiencing homelessness on their risk of harm if unhoused, to rank them for housing. Black clients were still served less often.
1 cited claim · 9 levers · Details · Open in the Lab
VI-SPDAT
The standard nobody validated: a homelessness triage score
The VI-SPDAT questionnaire ranked people experiencing homelessness for housing. It spread to dozens of U.S. states before anyone tested it. Its co-creator, OrgCode, withdrew it.
2 cited claims · 10 levers · Details · Open in the Lab
LA County Homelessness Prevention Unit
The help you have to be found for: a homelessness-prevention model
LA County offers cash and help to residents its model ranks at highest risk of homelessness. The model misses most who later become homeless.
2 cited claims · 10 levers · Details · Open in the Lab
Xantura OneView (predictive homelessness flagging)
The flag no one can reach: a predictive homelessness-prevention platform
Xantura's OneView combines council data to flag households at risk of homelessness. In Maidstone's pilot, one officer could follow up about 260 of 650-plus alerts.
2 cited claims · 11 levers · Details · Open in the Lab
CHAI (chronic-homelessness prediction)
The people the data can't see: a consent-based homelessness-risk model
CHAI tells London, Ontario caseworkers which shelter clients may become chronically homeless. It explains each flag and allows opt-outs, but sees only public-shelter users.
2 cited claims · 14 levers · Details · Open in the Lab
Imagine LA Benefit Navigator copilot
Best where you can check it least: a benefits-navigation copilot
A chatbot answers Los Angeles caseworkers' benefits questions, quoting policy. Accuracy rose most for new staff and hard questions, where errors are hardest to spot.
1 cited claim · 10 levers · Details · Open in the Lab
London's Strategic Insights Tool
One shared memory and thirty-three readers: consolidating a city's rough-sleeping records
London's Strategic Insights Tool links three record systems into one picture of people sleeping rough. All 33 local authorities read it to plan services.
2 cited claims · 10 levers · Details · Open in the Lab
LA's coordinated-entry triage revision
Two scores in one queue: the transition that let a retired bias back in
Los Angeles swapped a biased housing survey for a fairer score. LAHSA says clients qualified more easily on the old one, so providers kept it.
2 cited claims · 10 levers · Details · Open in the Lab
Santa Clara County Homelessness Prevention System
A measured lever on an unmeasured target: a homelessness-prevention screen
Santa Clara's questionnaire scores households seeking emergency money to keep their homes. A trial shows the money works. Does it go to the right households?
2 cited claims · 11 levers · Details · Open in the Lab
Calgary Drop-In Centre
The canvas rather than the answer: interpretable screening a shelter's own staff choose to check
At the Calgary Drop-In Centre, a homeless shelter, staff read client histories, not a score. They lean on data more for housing than for bans.
2 cited claims · 10 levers · Details · Open in the Lab
Homebase Risk Assessment Questionnaire
The prevention screener that has to describe itself in public every year
Caseworkers in Homebase, New York City's homelessness-prevention program, ask households fifteen questions. Based on the points total, they offer full services or a brief contact.
3 cited claims · 10 levers · Details · Open in the Lab
Immigration & asylum AI
2 examples
BAMF dialect recognition
One clue among many or the thing that decides
Germany's asylum agency uses DIAS, dialect-recognition software, to estimate applicants' origin from their speech. Does its estimate count for more than its accuracy supports?
2 cited claims · 11 levers · Details · Open in the Lab
Home Office IPIC
A human decides and the form only asks why not
IPIC, a Home Office algorithm, recommends migrants for immigration decisions or enforcement. Officials must justify rejecting its recommendation, not accepting it. Is their review real?
2 cited claims · 8 levers · Details · Open in the Lab
Industrial QA & operations AI
5 examples
Predictive maintenance on a high-speed rail fleet
The contract that prices every miss
Siemens analytics read sensor data from Renfe's Velaro E trains to forecast part failures. Renfe promises a full refund for delays over 15 minutes.
2 cited claims · 7 levers · Details · Open in the Lab
Audi press-shop inspection
The inspection the model inherited
Audi's software looks for hairline cracks in pressed sheet-metal parts. Audi said it would replace the earlier check by people and cameras.
2 cited claims · 9 levers · Details · Open in the Lab
BMW AIQX inspection
The flag is not the catch until someone acts on it
BMW's AIQX system flags defects to line workers as vehicles are assembled. A flag becomes a caught defect only when a worker checks it.
2 cited claims · 8 levers · Details · Open in the Lab
A heavy-industry predictive-maintenance deployment
Ninety percent fewer false alarms while the crews still label
In a peer-reviewed heavy-industry study, crews labeled a sensor model's alarms, and false alarms fell about 90 percent. The gain lasts while crews stay engaged.
2 cited claims · 8 levers · Details · Open in the Lab
Automated visual inspection of injectable drugs
Erring toward the scrap heap while guarding the one that gets through
Machine learning inspects filled injectable drugs for defects, tuned to err toward scrapping good vials rather than passing a bad one. Inspectors remain the backstop.
2 cited claims · 9 levers · Details · Open in the Lab
Lending & credit collections AI
4 examples
Upstart lending model
Regulator-verified access — and a search left at an impasse
Upstart's model approves, declines, and prices personal loans with no human review. Published reports show wider access to credit, and approval gaps for Black applicants.
2 cited claims · 5 levers · Details · Open in the Lab
Apple Card underwriting
Cleared on the numbers but unable to say why
In 2019, people complained the Apple Card gave women lower credit limits. New York's regulator found no unlawful discrimination but faulted customer service and transparency.
2 cited claims · 7 levers · Details · Open in the Lab
Earnest AI underwriting
A neutral-looking feature and the testing no one ran
Earnest's student loan models priced loans by a school's default rate and denied some non-citizens outright. A 2025 settlement bars both and requires fairness testing.
2 cited claims · 7 levers · Details · Open in the Lab
TransUnion OFAC Name Screen
Two fields, and the file that held the rest
TransUnion sold a credit report add-on that checked only a consumer's name against a Treasury sanctions list. An appeals court described thousands of false matches.
5 cited claims · 11 levers · Details · Open in the Lab
Logistics dispatch & scheduling AI
2 examples
UPS delivery route optimization
A hundred million miles saved and the discretion it cost
UPS's ORION system sets each driver's delivery route, saving miles and fuel. Vehicle data show whether drivers follow it, so the saving comes with monitoring.
2 cited claims · 9 levers · Details · Open in the Lab
Amazon fulfillment-centre algorithmic management
Units per hour up and a cost measured in bodies
Amazon's warehouse software sets workers' tasks and pace. Throughput rose, and a federal regulator and a Senate committee tied that pace to worker injuries.
2 cited claims · 10 levers · Details · Open in the Lab
Public benefits & eligibility
23 examples
Michigan MiDAS
Automation without review: a benefits-fraud system
Michigan's MiDAS system wrote unemployment fraud determinations into claimant records, often with no human review. It wrongly accused tens of thousands of people.
3 cited claims · 12 levers · Details · Open in the Lab
Michigan MiDAS
After the settlements: review returns to a benefits-fraud system
Michigan's MiDAS software decided many unemployment fraud cases without human review. A 2017 lawsuit settlement made review a requirement. This case asks what sustains it.
3 cited claims · 4 levers · Details · Open in the Lab
Rotterdam welfare-fraud risk model
The suspicion machine: a welfare-fraud risk model
Rotterdam's fraud risk model ranked welfare recipients for investigation. The city paused it after an audit, and journalists later documented its skew against vulnerable groups.
2 cited claims · 13 levers · Details · Open in the Lab
Arkansas ARChoices / ARIA
When the tool sets the hours: a home-care hours allocator
Arkansas's Medicaid program let an algorithm set disabled and older people's weekly home-care hours from a scored assessment. In 2016, nearly half had hours cut.
2 cited claims · 11 levers · Details · Open in the Lab
Netherlands childcare-benefits scandal (Toeslagenaffaire)
The institutional amplifier: a childcare-benefits fraud-hunt
The Dutch Tax Administration's benefits branch wrongly accused an estimated 26,000 or more families of childcare benefit fraud. It demanded they repay their whole allowance.
2 cited claims · 13 levers · Details · Open in the Lab
SyRI (Netherlands)
Struck down before the harm was counted: a secret welfare-fraud dragnet
SyRI, a secret Dutch system, linked government records to flag people for fraud investigation. In 2020 a court stopped it on privacy and transparency grounds.
2 cited claims · 11 levers · Details · Open in the Lab
CNAF benefit-fraud risk score (France)
The score that suspects the vulnerable: a benefit-fraud risk model
France's family-benefits fund, CNAF, scores every benefit-receiving household for fraud risk each month. In the versions examined, markers of economic vulnerability raised the score.
2 cited claims · 11 levers · Details · Open in the Lab
Forsakringskassan VAB fraud-selection profile (Sweden)
The audit the agency refused: a secret fraud-selection profile no one outside could see
Försäkringskassan, Sweden's Social Insurance Agency, used a machine-learning profile to pick sick-child benefit claimants for fraud and error investigation. It never released the model.
2 cited claims · 10 levers · Details · Open in the Lab
Udbetaling Danmark data-driven control (Denmark)
Standing surveillance by data-linking: a welfare-fraud control suite
Udbetaling Danmark's fraud-control models score Danish benefit recipients on data from about ten linked national registers. A control team decides which flagged cases to investigate.
2 cited claims · 12 levers · Details · Open in the Lab
BOSCO (Spain)
The secret code: an eligibility engine that gives no reasons
Spain's BOSCO software decides who gets an electricity-bill discount. It gives no reasons, and one flaw in its rules can deny thousands of eligible people.
2 cited claims · 9 levers · Details · Open in the Lab
Serbia Social Card (Socijalna karta)
Cut off by a data match: a social-assistance registry
Serbia's Social Card registry matches records to flag suspected income or assets. Its flags could cut assistance, were rarely contested, and were hard to correct.
2 cited claims · 10 levers · Details · Open in the Lab
UK DWP Universal Credit Advances fraud model
The self-audited skew: a benefits fraud-scoring model
A model scores Universal Credit advances for fraud risk. Its department published that it refers older and non-UK claimants more often, and kept it running.
2 cited claims · 11 levers · Details · Open in the Lab
ID.me identity verification
The gate nobody counts: an identity check in front of benefits
ID.me's facial-recognition check stood in front of pandemic unemployment claims in 25-plus states. A claimant who did not finish it was never recorded as denied.
2 cited claims · 10 levers · Details · Open in the Lab
Medicaid unwinding ex-parte renewals
The unit of determination: an automated renewal system at population scale
In 2023, 30 states' Medicaid renewal systems judged eligibility per household, not per person. States improperly ended coverage for nearly 500,000 people, many children.
1 cited claim · 10 levers · Details · Open in the Lab
INSS automated benefit analysis
Automation as queue management: when the metric makes denial the fastest way out
Brazil's social security institute, INSS, uses automation to clear a benefit backlog. It counts staff output in cases analyzed, and denial finishes a case fastest.
1 cited claim · 11 levers · Details · Open in the Lab
Samagra Vedika
The match that cancels you: entity resolution as eligibility
In India's Telangana state, Samagra Vedika matches people across thirty-plus databases. A similarly-named stranger's car, matched to a household, could cancel its ration card unannounced.
3 cited claims · 11 levers · Details · Open in the Lab
Workforce Australia Targeted Compliance Framework
The lesson not learned: automated compliance sanctioning after a scandal
From 2022, Australia's automated Targeted Compliance Framework cancelled payments without weighing jobseekers' reasonable excuses for missed requirements. The Ombudsman found 1,009 jobseekers' payments unlawfully…
1 cited claim · 11 levers · Details · Open in the Lab
NYC MyCity business chatbot
Exposure is not correction: a public-facing government advice chatbot
New York City's MyCity chatbot said business owners could break laws protecting workers and tenants. It stayed online roughly two years after reporters exposed it.
3 cited claims · 9 levers · Details · Open in the Lab
Nevada DETR generative-AI unemployment appeals
The referee who signs: an AI that drafts the ruling
Nevada's unemployment agency had Google build an AI that drafts appeal rulings for referees to sign. This case asks whether signing stays a review.
2 cited claims · 12 levers · Details · Open in the Lab
Tennessee TennCare TEDS
The notice that never came: an automated Medicaid eligibility system
TEDS decides Tennessee Medicaid eligibility and generates the notices people need to appeal. A federal court held its wrong terminations and misleading notices unlawful.
1 cited claim · 12 levers · Details · Open in the Lab
Robodebt (Australia)
The debt you must disprove: an income-averaging engine
Australia's Robodebt scheme raised welfare debts by spreading a person's yearly tax-office income evenly across fortnights. Recipients then had to disprove the debts.
2 cited claims · 10 levers · Details · Open in the Lab
Robodebt (Australia)
After the Commission: refunding the debts an engine raised
Robodebt raised hundreds of thousands of wrongful welfare debts in Australia. Outside controls ended it. This case asks what the agency needs to check itself.
2 cited claims · 5 levers · Details · Open in the Lab
Indiana / IBM eligibility modernization
Denied for 'failure to cooperate': a privatized eligibility pipeline
Indiana outsourced welfare eligibility to an IBM-led consortium. Over a million denials followed in its first years, many for procedural 'failure to cooperate'.
2 cited claims · 11 levers · Details · Open in the Lab
Security operations & fraud detection
3 examples
ML anti-money-laundering as primary monitoring
Fewer alerts and more confirmed — but confirmed by whom?
HSBC replaced rules-based anti-money-laundering monitoring with Google Cloud's AML AI. It reports more confirmed suspicious activity from fewer alerts, figures no one has independently audited.
2 cited claims · 9 levers · Details · Open in the Lab
Fraud false positives that froze real accounts
The wrong flag that took ninety days to reverse
Chime's fraud algorithms wrongly flagged legitimate customers, whose accounts were frozen or closed. A 2024 federal consent order penalized the delayed refunds, not the flags.
2 cited claims · 10 levers · Details · Open in the Lab
Danske Bank fraud scoring
Better detection but worse reimbursement — and a rule that moved it
Teradata's case study claims Danske Bank's fraud engine catches more fraud. Yet the bank later ranked worst among UK banks at reimbursing scam victims.
2 cited claims · 8 levers · Details · Open in the Lab
Software engineering AI (coding assistants)
4 examples
A commercial code assistant across three enterprises
Big lift for novices but slower for experts: a coding assistant
GitHub Copilot suggests code to developers. Trials at three companies found large gains for less-experienced developers. An independent study found experienced developers slower using AI.
2 cited claims · 11 levers · Details · Open in the Lab
Google ML code completion
Owning every node: the strength and the missing check
Google's platform team built a code-completion system for more than 10,000 Google developers, then measured it themselves. No outside party checked the results.
2 cited claims · 11 levers · Details · Open in the Lab
Gated coding-assistant rollout at a regulated bank
The gate that recorded what it couldn't resolve: a bank's rollout
ANZ Bank tried GitHub Copilot with about 100 engineers, then extended it to about 1,000. It recorded the tool's effect on security as inconclusive.
2 cited claims · 7 levers · Details · Open in the Lab
GitHub Copilot at ZoomInfo
Measured carefully but measuring the wrong thing: an ordinary rollout
ZoomInfo rolled out GitHub Copilot to over 400 developers in four phases. It measured acceptance and satisfaction, not delivered output, and reported no security evaluation.
2 cited claims · 8 levers · Details · Open in the Lab