Skip to content

Lab Index

Every example in the PAN Lab

127 organizational networks you can explore, stress test, and govern in PAN Lab v0.1. Each has its own page: what it models, the sources and evidence behind it, and the concepts, pressures and levers it connects to across the Governance Center.

Every context is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Some of these organizations depend on the same outside supplier, reviewer or assessment instrument as another one in this list. Common Cause reads across the index for those shared dependencies, and quotes the evidence for each one.

Every pressure the Lab can apply has a page of its own, with what it pushes on and which levers answer it. Read about the Lab's pressures

Every lever the Lab offers has a page of its own, with what it pushes on and which pressures it answers. Read about the Lab's levers

Compare approaches

Teaching networks that hold the model fixed and vary only governance — not models of one real deployment.

Behavioral-health & crisis triage

10 examples

  • REACH VET

    The flag that works and the outcome it misses: a suicide-risk model

    Each month REACH VET, a Veterans Affairs model, flags high-risk veterans for suicide-prevention outreach. Appointments rose, but two evaluations found no drop in suicide deaths.

    2 cited claims · 12 levers · Details · Open in the Lab

  • Vanderbilt VSAIL suicide-risk alert

    The alert that had to be dismissed: an EHR suicide-risk model

    In a trial, Vanderbilt's suicide-risk alert led clinicians to screen in 42 percent of flagged visits as a pop-up, 4 percent as a chart icon.

    2 cited claims · 11 levers · Details · Open in the Lab

  • Kaiser Permanente Suicide-Risk Model

    The added sensor: an EHR-embedded suicide-risk score

    Kaiser Permanente's machine-learning score flags intake patients at risk of a suicide attempt. The flag, or a positive questionnaire, prompts a clinician to assess them.

    2 cited claims · 9 levers · Details · Open in the Lab

  • Crisis Text Line & Loris.ai

    The corpus and the spinoff: governing crisis-conversation data

    Crisis Text Line, a crisis texting service, ranks waiting texters with a model. Loris.ai, a for-profit it partly owned, trained software on their conversations.

    2 cited claims · 10 levers · Details · Open in the Lab

  • NarxCare

    The score you cannot see: an opaque prescribing-risk model

    Bamboo Health's NarxCare turns state prescription records into patient risk scores. Patients cannot see or contest the scores, which can gate access to pain medication.

    2 cited claims · 13 levers · Details · Open in the Lab

  • Limbic Access (NHS Talking Therapies)

    The front door that works: a self-referral triage chatbot

    The Limbic Access chatbot takes NHS talking-therapy self-referrals. Referrals rose more in services using it. Nearly all evidence it works comes from its maker.

    2 cited claims · 9 levers · Details · Open in the Lab

  • Woebot (a governed app wind-down)

    The responsible wind-down: retiring a peer-reviewed CBT chatbot

    In 2025 Woebot Health retired Woebot, its therapy chatbot app, on a published schedule. Users could request their chat transcripts before account data was anonymized.

    2 cited claims · 9 levers · Details · Open in the Lab

  • LyssnCrisis counselor QA at ProtoCall Services (988)

    The AI watches the counselor rather than the caller: a crisis-line QA scorer

    LyssnCrisis rates how counselors on 988, the US crisis line, handle calls, never the callers. No outside study has checked its scores or their benefit.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Gaggle Safety Management

    The discontinuation that changed the vendor rather than the graph: a school monitoring scanner

    Gaggle's software scans students' school accounts for signs of self-harm. When students sued, one district switched vendors. Did the practice change, or only the name?

    2 cited claims · 11 levers · Details · Open in the Lab

  • Oxevision camera monitoring on NHS mental health wards

    The evaluation was written by the seller: a bedroom monitor no one independent checked

    NHS mental health wards film patients' bedrooms with Oxevision cameras. One hospital trust's procedure set consent aside, and the seller shaped the evidence for it.

    2 cited claims · 10 levers · Details · Open in the Lab

Benefits navigation & public-facing chat

11 examples

  • Nava assistive benefits chatbot

    Done carefully: a verify-before-use copilot

    Nava's chatbot answers caseworkers' benefits questions from vetted documents, with quotes to check. The question is what keeps this cautious design in place under pressure.

    10 levers · Details · Open in the Lab

  • GOV.UK Chat

    The gate that said not yet: a public assistant behind a staged pilot gate

    GOV.UK Chat answers the public's questions from official guidance. Its gate held back a version that fell short, but its builders grade its accuracy themselves.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Burokratt

    The network of networks: a federated public-service chatbot

    Estonia's Burokratt puts many public institutions' chatbots behind one window. A mistake in its shared routing or knowledge can appear in every institution's answers.

    1 cited claim · 11 levers · Details · Open in the Lab

  • Caddy adviser copilot at Citizens Advice

    The gate is a job rather than a habit: a supervisor-checked adviser copilot

    Caddy drafts benefits answers for Citizens Advice advisers, and a supervisor checks every draft before an adviser sees it. Can that check survive wider use?

    2 cited claims · 10 levers · Details · Open in the Lab

  • Albert France Services

    Killed without a number: a sovereign adviser assistant that no metric ever measured

    Albert, a French state AI assistant, drafted answers for France Services advisers in a pilot. No error rate was published, and it was not extended.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Benefits Data Trust wind-down

    The node that could not be kept: winding down a benefits-navigation nonprofit

    In June 2024 Benefits Data Trust's board voted to close the benefits-navigation nonprofit within 60 days. Its government partners had no designated successor.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Propel in-app SNAP benefits assistant

    Read-only by design: a benefits assistant grounded on the record it never writes

    Propel's app uses AI to help people learn why a SNAP food-benefit deposit did not arrive. It reads the state's record, never writing to it.

    2 cited claims · 9 levers · Details · Open in the Lab

  • GetCalFresh

    Most of a state's online intake with no authority at all: an assisted-application node and its handoff

    GetCalFresh, a nonprofit's online form, carried most of California's online food-benefit applications but decided none. What if a service no law required carries that much?

    2 cited claims · 10 levers · Details · Open in the Lab

  • Frida (NAV Norway)

    Ask for a human: the handover boundary as a governed surface

    Frida is the chatbot of NAV, Norway's welfare agency. The case asks how easily citizens reach a person and who reads chats ending without one.

    2 cited claims · 11 levers · Details · Open in the Lab

  • IRS collection chatbots

    Expanded without a ruler: a federal collection chatbot with no performance measures

    The Internal Revenue Service (IRS) expanded its collection chatbot and live chat, and made live chat permanent, with no performance measures or reliable statistics.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Singapore's chatbot fleet refresh

    Eighty engines into one: a whole-of-government chatbot fleet refresh

    Singapore set out to move its agencies' chatbots off scripted engines, which match keywords to prewritten answers, onto a few shared AI engines.

    2 cited claims · 10 levers · Details · Open in the Lab

Caseworker documentation & copilots

8 examples

  • Magic Notes (Beam)

    The drafted record: a case-notes copilot

    UK council social care staff draft case notes, the client's record, with Magic Notes, AI software from the company Beam. They review each draft first.

    10 levers · Details · Open in the Lab

  • Guided walkthrough: agentic low-oversight office

    A walkthrough-only version of the agentic low-oversight office: an AI assistant that also acts as an autonomous agent, a stretched staff supervising many automated actions at once, and — added for the tour — the Outside…

    1 cited claim · 12 levers · Details · Open in the Lab

  • Minute / Local Transcribe

    The governed thing and the measured thing: a state-built meeting scribe

    The UK government built Minute, an AI meeting scribe, and piloted it in local councils. Its approval checks covered data handling and process, not accuracy.

    1 cited claim · 10 levers · Details · Open in the Lab

  • DWP Whitemail Insights and Vulnerability Scanner

    The letter no one reads twice: an upstream vulnerability scanner

    The UK welfare department's AI scans about 25,000 letters daily, flagging signs of risk like self-harm for caseworkers. No one rechecks the rest for risk.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Justice Transcribe

    The note that scores you: a copilot at the head of a risk pipeline

    Justice Transcribe drafts summaries of probation meetings for the case record in England and Wales. The ministry's reoffending-risk tools draw on probation records too.

    1 cited claim · 12 levers · Details · Open in the Lab

  • VA claims automation (automated survivor-benefit decisions)

    The record no one read: automated survivor-benefit decisions

    The VA's automation grants survivor benefits with no human involvement when its rules match. A 2026 audit found deficiencies in nearly all such grants.

    2 cited claims · 11 levers · Details · Open in the Lab

  • Amsterdam Smart Check

    The governed exit: a fair welfare screener that shipped every safeguard but one

    Slimme Check flagged welfare applications for investigation. Its bias, reduced on past data, shifted to other groups in a live pilot. Amsterdam halted it.

    1 cited claim · 12 levers · Details · Open in the Lab

  • Massachusetts DTA call summaries

    The record and its source: AI summaries of benefits calls

    Massachusetts piloted AI summaries of food-benefit eligibility calls. Each summary is saved into the case record, and the full call transcript is not kept.

    5 cited claims · 8 levers · Details · Open in the Lab

Child welfare & family services

15 examples

  • Allegheny Family Screening Tool

    The score and the screener: a child-welfare risk tool

    The Allegheny Family Screening Tool scores reports to the county's child protection hotline. Hotline staff weigh the score when deciding which families to investigate.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Illinois Rapid Safety Feedback

    The alarm that cried wolf: a child-safety scorer

    Illinois's child-welfare agency scored children in reports of suspected abuse or neglect. Its tool flagged thousands as extreme risks and missed children who died.

    2 cited claims · 13 levers · Details · Open in the Lab

  • Allegheny Hello Baby

    The help that keeps a file: a birth-risk prevention model

    Allegheny County's Hello Baby model scores newborns from county records. Highest-scored families are offered voluntary help, not investigation, by workers who must report suspected abuse.

    1 cited claim · 10 levers · Details · Open in the Lab

  • Douglas County Decision Aide

    The score read only at the edges: a child-welfare screening aide

    Douglas County, Colorado's Decision Aide scores reports of possible child maltreatment (referrals). An independent trial found it sped decisions without significantly changing children's outcomes.

    1 cited claim · 10 levers · Details · Open in the Lab

  • Eckerd Rapid Safety Feedback

    Endorsed then evaluated: a risk tool that spread on a claim

    Eckerd, a child-welfare nonprofit, built Rapid Safety Feedback to flag cases resembling past child deaths. It spread to several states before anyone independent tested it.

    2 cited claims · 11 levers · Details · Open in the Lab

  • ProKid (Netherlands)

    The colour and the record: a child risk-profiler

    ProKid, a Dutch police instrument, sorted children under 12 into colour risk bands from police records, including children recorded only as victims or witnesses.

    1 cited claim · 9 levers · Details · Open in the Lab

  • Insight Bristol / Think Family Database

    The database nobody could audit: a shared child-risk profiling system

    Bristol's council and police scored children from a shared database. Staff distrusted the scores. Two models were withdrawn, reportedly with no record of why.

    2 cited claims · 12 levers · Details · Open in the Lab

  • Sistema Alerta Niñez (Chile)

    Scored before anyone knocks: a child-risk targeting tool

    Chile's Sistema Alerta Niñez ranks children by risk, from data families gave for benefits. Local offices use it to decide who gets early help first.

    1 cited claim · 11 levers · Details · Open in the Lab

  • Los Angeles County Project AURA

    Caught at the gate: a child-abuse risk model that never shipped

    Los Angeles County tested Project AURA, a child-fatality risk score, on past cases. It falsely flagged 3,829 children, correctly flagged 171, and was never used.

    1 cited claim · 12 levers · Details · Open in the Lab

  • What Works for Children's Social Care ML pilots

    The bar it never cleared: a child-welfare prediction pilot

    A research centre's 32 models predicted which children's cases would escalate. None was right often enough to clear its published bar, so none went live.

    1 cited claim · 11 levers · Details · Open in the Lab

  • New Zealand MSD Predictive Risk Modelling

    Halted before it ran: a national child-risk model

    New Zealand's Ministry of Social Development commissioned a model to score newborns' maltreatment risk. A minister halted its study, and the tool never went live.

    1 cited claim · 11 levers · Details · Open in the Lab

  • Gladsaxe model

    The screen that never ran: whole-population child scoring

    Gladsaxe, a Danish municipality, built a model to score every young child for vulnerability. It was stopped in development and never ran on live cases.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Hackney / Xantura Early Help Profiling

    The pilot that quietly failed: a small council's family profiler

    Hackney Council paid Xantura to flag at-risk families, without telling them. Poor data meant the pilot found few new families, so the council dropped it.

    3 cited claims · 11 levers · Details · Open in the Lab

  • Oregon Safety at Screening

    The fix then the off switch: a fairness-corrected screening tool

    Oregon's child-welfare hotline screeners saw a fairness-corrected risk score before deciding whether to investigate a report. The agency stopped using the tool in 2022.

    1 cited claim · 12 levers · Details · Open in the Lab

  • Predict-Align-Prevent

    The map, not the score: a place-based risk surface and the records it concentrates

    A nonprofit's model ranks small map cells by child maltreatment risk, scoring no family. In New Hampshire, teams aimed outreach and funding where it pointed.

    3 cited claims · 13 levers · Details · Open in the Lab

Clinical decision support & deterioration alerting

6 examples

  • TREWS sepsis early-warning system

    The alert that works only when confirmed: a sepsis early-warning model

    TREWS, built at Johns Hopkins, scores inpatients at five Hopkins hospitals for sepsis. Its benefit came only from alerts a provider confirmed within three hours.

    2 cited claims · 8 levers · Details · Open in the Lab

  • Advance Alert Monitor (AAM) deterioration model

    The alert that never reaches the bedside: a screened deterioration model

    Kaiser Permanente's Advance Alert Monitor flags hospital patients predicted to worsen in about twelve hours. Regional nurses screen every alert before the bedside team acts.

    1 cited claim · 8 levers · Details · Open in the Lab

  • Sepsis Watch deep-learning detection system

    The nurse who gets the alert can't give the order: an authority-split detector

    Duke University Hospital's Sepsis Watch alerts a rapid response nurse, who cannot order treatment. The alert becomes care only if the nurse persuades a physician.

    2 cited claims · 8 levers · Details · Open in the Lab

  • Epic Sepsis Model

    Switched on before anyone checked: a proprietary sepsis model at scale

    Hundreds of hospitals switched on the Epic Sepsis Model before anyone independent checked it. When researchers did, it caught about a third of sepsis cases.

    2 cited claims · 8 levers · Details · Open in the Lab

  • nH Predict Utilization Review

    The order of operations inside a coverage determination

    UnitedHealthcare and naviHealth, both owned by UnitedHealth Group, decide whether members leaving the hospital get skilled nursing or rehabilitation care covered, and for how long.

    4 cited claims · 14 levers · Details · Open in the Lab

  • CA-CDS Child Abuse Alerting

    The mandated report: alerting that writes into another organisation

    A consortium at UPMC Children's Hospital of Pittsburgh built child-abuse alerts for emergency departments. Clinicians who suspect abuse must report it to a state agency.

    4 cited claims · 9 levers · Details · Open in the Lab

Clinical documentation copilots (ambient scribes)

3 examples

Content moderation & editorial AI

6 examples

  • YouTube Covid-19 enforcement

    The reviewers went home and the error rate showed

    With reviewers home in the pandemic, YouTube chose to over-remove by machine. Removals more than doubled. About half of appeals succeeded, up from a quarter.

    2 cited claims · 8 levers · Details · Open in the Lab

  • Meta content enforcement

    Ninety percent overturned on the cases chosen to be seen

    Meta's software makes millions of content decisions. The Oversight Board picks a few, some likely to be wrong, and in 2023 overturned about 90 percent.

    2 cited claims · 8 levers · Details · Open in the Lab

  • Sports Illustrated AI bylines

    The byline nobody was behind

    Sports Illustrated published product reviews by invented writers, produced by its contractor AdVon Commerce. An outside investigation found the fabrication, not the outlet's own editors.

    2 cited claims · 7 levers · Details · Open in the Lab

  • CNET AI-drafted articles

    Half the articles corrected under a byline that promised a review

    CNET published 77 AI-drafted finance explainers under a staff byline. Its own audit found that 41 needed correction. The byline implied a human review.

    2 cited claims · 8 levers · Details · Open in the Lab

  • StopNCII & Take It Down

    Two ways to run a hash bank: the intimate-image removal pair

    StopNCII and Take It Down fingerprint a person's intimate images on their own device, so partner platforms can find copies shared without consent.

    5 cited claims · 12 levers · Details · Open in the Lab

  • YouTube's Content ID copyright matching system

    The party that gains answers the objection

    YouTube's Content ID checks every upload against rights-holders' reference files. On a match, the rights-holder's preset instruction applies, and that rights-holder answers any dispute.

    8 cited claims · 10 levers · Details · Open in the Lab

Customer service & contact-centre AI

4 examples

  • A contact centre's generative-AI agent assist

    Fifteen percent on average and thirty for the newcomers

    A Fortune 500 software firm gave support agents an AI copilot. On average they resolved 15 percent more issues per hour. Newer agents gained most.

    2 cited claims · 7 levers · Details · Open in the Lab

  • DPD customer-support chatbot

    The guardrails that stopped holding after an update

    After a system update, a customer prompted DPD's support chatbot to swear and write a poem calling DPD the worst delivery firm in the world.

    2 cited claims · 7 levers · Details · Open in the Lab

  • Air Canada chatbot

    A policy the chatbot invented and the company that answered for it

    Air Canada's website chatbot told a customer they could claim the reduced bereavement fare after booking. The airline refused, and a tribunal held it liable.

    2 cited claims · 7 levers · Details · Open in the Lab

  • Klarna AI assistant

    Two-thirds of chats handled and a year later a rethink

    Klarna said its AI assistant handled about two-thirds of customer-service chats in the assistant's first month. About a year later, Klarna said quality had fallen.

    2 cited claims · 8 levers · Details · Open in the Lab

Education AI

3 examples

  • Chicago Public Schools On-Track indicator

    The rule a teacher can explain

    Chicago Public Schools uses a readable rule, not AI, to flag ninth graders as off track to graduate. Graduation rates later rose to record highs.

    2 cited claims · 7 levers · Details · Open in the Lab

  • Wisconsin DEWS

    Wrong most of the time and unevenly by race

    Wisconsin's Dropout Early Warning System labeled students 'high risk'. An audit found most graduated anyway, and Black and Hispanic students were wrongly labeled more often.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Cleveland State remote proctoring

    A camera in the bedroom and a flag that lands by skin tone

    Cleveland State University required a webcam scan of students' rooms before online exams, and proctoring software flagged suspected cheating. A court held the scan unconstitutional.

    2 cited claims · 8 levers · Details · Open in the Lab

Hiring & employment screening AI

4 examples

  • Amazon recruiting engine

    It learned who was hired rather than who succeeds

    Amazon's experimental resume scorer, trained on ten years of resumes mostly from men, penalized "women's". Amazon scrapped it when edits could not guarantee a fix.

    2 cited claims · 8 levers · Details · Open in the Lab

  • Workday AI screening

    One model across thousands of employers and testing no one can see

    Workday's AI screens applicants for thousands of employers. A lawsuit alleges age discrimination, and a court ruled Workday need not hand over its bias testing.

    2 cited claims · 8 levers · Details · Open in the Lab

  • Unilever and HireVue graduate hiring

    Good audits with real savings — and a cohort no one can see

    Unilever screened recent graduates with two vendors' AI tools: pymetrics games, then HireVue video-interview scoring. Its reported gains cover only the applicants it advanced.

    2 cited claims · 7 levers · Details · Open in the Lab

  • Intuit's recorded video assessment for promotion

    Assessing an incumbent: a recorded promotion gate and the captioning request

    Intuit put a Deaf employee's promotion through a recorded HireVue video assessment. A pending charge alleges Intuit denied the human captioning she asked for.

    7 cited claims · 12 levers · Details · Open in the Lab

Housing & homelessness services

11 examples

  • Allegheny Housing Assessment

    The score and the scarce bed: a coordinated-entry housing tool

    Allegheny County scores people experiencing homelessness on their risk of harm if unhoused, to rank them for housing. Black clients were still served less often.

    1 cited claim · 9 levers · Details · Open in the Lab

  • VI-SPDAT

    The standard nobody validated: a homelessness triage score

    The VI-SPDAT questionnaire ranked people experiencing homelessness for housing. It spread to dozens of U.S. states before anyone tested it. Its co-creator, OrgCode, withdrew it.

    2 cited claims · 10 levers · Details · Open in the Lab

  • LA County Homelessness Prevention Unit

    The help you have to be found for: a homelessness-prevention model

    LA County offers cash and help to residents its model ranks at highest risk of homelessness. The model misses most who later become homeless.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Xantura OneView (predictive homelessness flagging)

    The flag no one can reach: a predictive homelessness-prevention platform

    Xantura's OneView combines council data to flag households at risk of homelessness. In Maidstone's pilot, one officer could follow up about 260 of 650-plus alerts.

    2 cited claims · 11 levers · Details · Open in the Lab

  • CHAI (chronic-homelessness prediction)

    The people the data can't see: a consent-based homelessness-risk model

    CHAI tells London, Ontario caseworkers which shelter clients may become chronically homeless. It explains each flag and allows opt-outs, but sees only public-shelter users.

    2 cited claims · 14 levers · Details · Open in the Lab

  • Imagine LA Benefit Navigator copilot

    Best where you can check it least: a benefits-navigation copilot

    A chatbot answers Los Angeles caseworkers' benefits questions, quoting policy. Accuracy rose most for new staff and hard questions, where errors are hardest to spot.

    1 cited claim · 10 levers · Details · Open in the Lab

  • London's Strategic Insights Tool

    One shared memory and thirty-three readers: consolidating a city's rough-sleeping records

    London's Strategic Insights Tool links three record systems into one picture of people sleeping rough. All 33 local authorities read it to plan services.

    2 cited claims · 10 levers · Details · Open in the Lab

  • LA's coordinated-entry triage revision

    Two scores in one queue: the transition that let a retired bias back in

    Los Angeles swapped a biased housing survey for a fairer score. LAHSA says clients qualified more easily on the old one, so providers kept it.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Santa Clara County Homelessness Prevention System

    A measured lever on an unmeasured target: a homelessness-prevention screen

    Santa Clara's questionnaire scores households seeking emergency money to keep their homes. A trial shows the money works. Does it go to the right households?

    2 cited claims · 11 levers · Details · Open in the Lab

  • Calgary Drop-In Centre

    The canvas rather than the answer: interpretable screening a shelter's own staff choose to check

    At the Calgary Drop-In Centre, a homeless shelter, staff read client histories, not a score. They lean on data more for housing than for bans.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Homebase Risk Assessment Questionnaire

    The prevention screener that has to describe itself in public every year

    Caseworkers in Homebase, New York City's homelessness-prevention program, ask households fifteen questions. Based on the points total, they offer full services or a brief contact.

    3 cited claims · 10 levers · Details · Open in the Lab

Immigration & asylum AI

2 examples

  • BAMF dialect recognition

    One clue among many or the thing that decides

    Germany's asylum agency uses DIAS, dialect-recognition software, to estimate applicants' origin from their speech. Does its estimate count for more than its accuracy supports?

    2 cited claims · 11 levers · Details · Open in the Lab

  • Home Office IPIC

    A human decides and the form only asks why not

    IPIC, a Home Office algorithm, recommends migrants for immigration decisions or enforcement. Officials must justify rejecting its recommendation, not accepting it. Is their review real?

    2 cited claims · 8 levers · Details · Open in the Lab

Industrial QA & operations AI

5 examples

Lending & credit collections AI

4 examples

  • Upstart lending model

    Regulator-verified access — and a search left at an impasse

    Upstart's model approves, declines, and prices personal loans with no human review. Published reports show wider access to credit, and approval gaps for Black applicants.

    2 cited claims · 5 levers · Details · Open in the Lab

  • Apple Card underwriting

    Cleared on the numbers but unable to say why

    In 2019, people complained the Apple Card gave women lower credit limits. New York's regulator found no unlawful discrimination but faulted customer service and transparency.

    2 cited claims · 7 levers · Details · Open in the Lab

  • Earnest AI underwriting

    A neutral-looking feature and the testing no one ran

    Earnest's student loan models priced loans by a school's default rate and denied some non-citizens outright. A 2025 settlement bars both and requires fairness testing.

    2 cited claims · 7 levers · Details · Open in the Lab

  • TransUnion OFAC Name Screen

    Two fields, and the file that held the rest

    TransUnion sold a credit report add-on that checked only a consumer's name against a Treasury sanctions list. An appeals court described thousands of false matches.

    5 cited claims · 11 levers · Details · Open in the Lab

Logistics dispatch & scheduling AI

2 examples

Public benefits & eligibility

23 examples

  • Michigan MiDAS

    Automation without review: a benefits-fraud system

    Michigan's MiDAS system wrote unemployment fraud determinations into claimant records, often with no human review. It wrongly accused tens of thousands of people.

    3 cited claims · 12 levers · Details · Open in the Lab

  • Michigan MiDAS

    After the settlements: review returns to a benefits-fraud system

    Michigan's MiDAS software decided many unemployment fraud cases without human review. A 2017 lawsuit settlement made review a requirement. This case asks what sustains it.

    3 cited claims · 4 levers · Details · Open in the Lab

  • Rotterdam welfare-fraud risk model

    The suspicion machine: a welfare-fraud risk model

    Rotterdam's fraud risk model ranked welfare recipients for investigation. The city paused it after an audit, and journalists later documented its skew against vulnerable groups.

    2 cited claims · 13 levers · Details · Open in the Lab

  • Arkansas ARChoices / ARIA

    When the tool sets the hours: a home-care hours allocator

    Arkansas's Medicaid program let an algorithm set disabled and older people's weekly home-care hours from a scored assessment. In 2016, nearly half had hours cut.

    2 cited claims · 11 levers · Details · Open in the Lab

  • Netherlands childcare-benefits scandal (Toeslagenaffaire)

    The institutional amplifier: a childcare-benefits fraud-hunt

    The Dutch Tax Administration's benefits branch wrongly accused an estimated 26,000 or more families of childcare benefit fraud. It demanded they repay their whole allowance.

    2 cited claims · 13 levers · Details · Open in the Lab

  • SyRI (Netherlands)

    Struck down before the harm was counted: a secret welfare-fraud dragnet

    SyRI, a secret Dutch system, linked government records to flag people for fraud investigation. In 2020 a court stopped it on privacy and transparency grounds.

    2 cited claims · 11 levers · Details · Open in the Lab

  • CNAF benefit-fraud risk score (France)

    The score that suspects the vulnerable: a benefit-fraud risk model

    France's family-benefits fund, CNAF, scores every benefit-receiving household for fraud risk each month. In the versions examined, markers of economic vulnerability raised the score.

    2 cited claims · 11 levers · Details · Open in the Lab

  • Forsakringskassan VAB fraud-selection profile (Sweden)

    The audit the agency refused: a secret fraud-selection profile no one outside could see

    Försäkringskassan, Sweden's Social Insurance Agency, used a machine-learning profile to pick sick-child benefit claimants for fraud and error investigation. It never released the model.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Udbetaling Danmark data-driven control (Denmark)

    Standing surveillance by data-linking: a welfare-fraud control suite

    Udbetaling Danmark's fraud-control models score Danish benefit recipients on data from about ten linked national registers. A control team decides which flagged cases to investigate.

    2 cited claims · 12 levers · Details · Open in the Lab

  • BOSCO (Spain)

    The secret code: an eligibility engine that gives no reasons

    Spain's BOSCO software decides who gets an electricity-bill discount. It gives no reasons, and one flaw in its rules can deny thousands of eligible people.

    2 cited claims · 9 levers · Details · Open in the Lab

  • Serbia Social Card (Socijalna karta)

    Cut off by a data match: a social-assistance registry

    Serbia's Social Card registry matches records to flag suspected income or assets. Its flags could cut assistance, were rarely contested, and were hard to correct.

    2 cited claims · 10 levers · Details · Open in the Lab

  • UK DWP Universal Credit Advances fraud model

    The self-audited skew: a benefits fraud-scoring model

    A model scores Universal Credit advances for fraud risk. Its department published that it refers older and non-UK claimants more often, and kept it running.

    2 cited claims · 11 levers · Details · Open in the Lab

  • ID.me identity verification

    The gate nobody counts: an identity check in front of benefits

    ID.me's facial-recognition check stood in front of pandemic unemployment claims in 25-plus states. A claimant who did not finish it was never recorded as denied.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Medicaid unwinding ex-parte renewals

    The unit of determination: an automated renewal system at population scale

    In 2023, 30 states' Medicaid renewal systems judged eligibility per household, not per person. States improperly ended coverage for nearly 500,000 people, many children.

    1 cited claim · 10 levers · Details · Open in the Lab

  • INSS automated benefit analysis

    Automation as queue management: when the metric makes denial the fastest way out

    Brazil's social security institute, INSS, uses automation to clear a benefit backlog. It counts staff output in cases analyzed, and denial finishes a case fastest.

    1 cited claim · 11 levers · Details · Open in the Lab

  • Samagra Vedika

    The match that cancels you: entity resolution as eligibility

    In India's Telangana state, Samagra Vedika matches people across thirty-plus databases. A similarly-named stranger's car, matched to a household, could cancel its ration card unannounced.

    3 cited claims · 11 levers · Details · Open in the Lab

  • Workforce Australia Targeted Compliance Framework

    The lesson not learned: automated compliance sanctioning after a scandal

    From 2022, Australia's automated Targeted Compliance Framework cancelled payments without weighing jobseekers' reasonable excuses for missed requirements. The Ombudsman found 1,009 jobseekers' payments unlawfully…

    1 cited claim · 11 levers · Details · Open in the Lab

  • NYC MyCity business chatbot

    Exposure is not correction: a public-facing government advice chatbot

    New York City's MyCity chatbot said business owners could break laws protecting workers and tenants. It stayed online roughly two years after reporters exposed it.

    3 cited claims · 9 levers · Details · Open in the Lab

  • Nevada DETR generative-AI unemployment appeals

    The referee who signs: an AI that drafts the ruling

    Nevada's unemployment agency had Google build an AI that drafts appeal rulings for referees to sign. This case asks whether signing stays a review.

    2 cited claims · 12 levers · Details · Open in the Lab

  • Tennessee TennCare TEDS

    The notice that never came: an automated Medicaid eligibility system

    TEDS decides Tennessee Medicaid eligibility and generates the notices people need to appeal. A federal court held its wrong terminations and misleading notices unlawful.

    1 cited claim · 12 levers · Details · Open in the Lab

  • Robodebt (Australia)

    The debt you must disprove: an income-averaging engine

    Australia's Robodebt scheme raised welfare debts by spreading a person's yearly tax-office income evenly across fortnights. Recipients then had to disprove the debts.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Robodebt (Australia)

    After the Commission: refunding the debts an engine raised

    Robodebt raised hundreds of thousands of wrongful welfare debts in Australia. Outside controls ended it. This case asks what the agency needs to check itself.

    2 cited claims · 5 levers · Details · Open in the Lab

  • Indiana / IBM eligibility modernization

    Denied for 'failure to cooperate': a privatized eligibility pipeline

    Indiana outsourced welfare eligibility to an IBM-led consortium. Over a million denials followed in its first years, many for procedural 'failure to cooperate'.

    2 cited claims · 11 levers · Details · Open in the Lab

Security operations & fraud detection

3 examples

  • ML anti-money-laundering as primary monitoring

    Fewer alerts and more confirmed — but confirmed by whom?

    HSBC replaced rules-based anti-money-laundering monitoring with Google Cloud's AML AI. It reports more confirmed suspicious activity from fewer alerts, figures no one has independently audited.

    2 cited claims · 9 levers · Details · Open in the Lab

  • Fraud false positives that froze real accounts

    The wrong flag that took ninety days to reverse

    Chime's fraud algorithms wrongly flagged legitimate customers, whose accounts were frozen or closed. A 2024 federal consent order penalized the delayed refunds, not the flags.

    2 cited claims · 10 levers · Details · Open in the Lab

  • Danske Bank fraud scoring

    Better detection but worse reimbursement — and a rule that moved it

    Teradata's case study claims Danske Bank's fraud engine catches more fraud. Yet the bank later ranked worst among UK banks at reimbursing scam victims.

    2 cited claims · 8 levers · Details · Open in the Lab

Software engineering AI (coding assistants)

4 examples

  • A commercial code assistant across three enterprises

    Big lift for novices but slower for experts: a coding assistant

    GitHub Copilot suggests code to developers. Trials at three companies found large gains for less-experienced developers. An independent study found experienced developers slower using AI.

    2 cited claims · 11 levers · Details · Open in the Lab

  • Google ML code completion

    Owning every node: the strength and the missing check

    Google's platform team built a code-completion system for more than 10,000 Google developers, then measured it themselves. No outside party checked the results.

    2 cited claims · 11 levers · Details · Open in the Lab

  • Gated coding-assistant rollout at a regulated bank

    The gate that recorded what it couldn't resolve: a bank's rollout

    ANZ Bank tried GitHub Copilot with about 100 engineers, then extended it to about 1,000. It recorded the tool's effect on security as inconclusive.

    2 cited claims · 7 levers · Details · Open in the Lab

  • GitHub Copilot at ZoomInfo

    Measured carefully but measuring the wrong thing: an ordinary rollout

    ZoomInfo rolled out GitHub Copilot to over 400 developers in four phases. It measured acceptance and satisfaction, not delivered output, and reported no security evaluation.

    2 cited claims · 8 levers · Details · Open in the Lab