PAN Lab example
Rotterdam welfare-fraud risk model
The suspicion machine: a welfare-fraud risk model
Rotterdam's fraud risk model ranked welfare recipients for investigation. The city paused it after an audit, and journalists later documented its skew against vulnerable groups.
See more
Rotterdam's welfare-fraud risk model was a machine-learning model, originally built with Accenture. The city of Rotterdam used it to score welfare recipients for fraud risk from about 315 inputs, including age, gender, language, and neighbourhood. Its ranked list set who fraud investigators looked at first.
How it was used
Investigators acted on the model's ranked list. The sources describe human review of the ranking as nominal: the ranking drove who was scrutinised.
Investigators recorded each investigation's findings in the recipient's dossier, and the model's risk score was logged there too. They read the dossier again when that person was investigated again.
The sources list human review and an appeal process for this deployment, but describe neither. From 2017 to 2021, officials ranked every benefit recipient in the city by score. Those in the top 10 percent were referred for investigation. The list was one of three ways the city chose whom to interview.
What the reviews found
In 2021 the city's own audit office, the Rekenkamer Rotterdam, published a report on the city's algorithms, Gekleurde technologie. It flagged this model's ethical risks, including proxy discrimination from inputs such as language skill. Proxy discrimination means an input stands in for a trait the law protects from discrimination, such as gender. The city put the model on hold in 2021.
In 2023, journalists at Lighthouse Reports and partners obtained the model and its training data. Their investigation, Suspicion Machines, documented how scores skewed against already vulnerable groups: women, parents, and people with limited Dutch. They found that people who had not passed the Dutch language requirement were almost twice as likely to be flagged.
WIRED and Follow the Money published their own accounts of the model. The Racism and Technology Center reported that it was biased.
Why this case matters
The case file calls Rotterdam its clearest case of an audit acting on a system. The city paused the model after its audit office's 2021 report. In 2023, journalists' access to the model itself documented the bias in detail. Suppliers usually keep their models closed to that kind of outside inspection.
The case also shows a harm that accuracy measures miss. The model decided who got investigated, and an investigation is itself a burden. Asking only whether a flag was correct misses the question of who has to go through an investigation.
The loop to watch
This network assumes a retraining loop. Past dossiers and the outcomes investigators recorded are used to train the next version of the model.
The model then learns from cases its own ranking chose to investigate. So a skew in who gets investigated can return in whom it suspects next.
The sources show the model had training data, which journalists obtained in 2023. Lighthouse Reports' published method says the city retrained the model each year with new data. Follow the Money reports it was trained on 12,700 recipients who had already been investigated.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. Here a mistake is, for example, a recipient ranked high risk because of their language or neighbourhood, or a wrong finding entered in a dossier. A pathway counts as closed once it passes on only a few mistakes. It need not stop them all.
This case has a budget of 15 units. Explore (No Targets) sets no targets. There, two tools costing 4 units are enough to stop mistakes building on one another across the network. Under Service Targets Only, the same two also meet the service target. One such pair is Mark AI-written records with Escalate checks. Here Mark AI-written records marks the model's scores in the dossiers, so investigators reading them, and the retraining, can weigh them.
Under Service and Safety Targets and All Governance Targets, the targets can also be met. Both levels ask you to close every failure pathway, among other targets. The cheapest combinations under Service and Safety Targets use five tools and cost 11 units. One is Mark AI-written records, Keep prompts neutral, Peer sharing rules, Escalate checks, and Store less data. This model takes no prompts: here Keep prompts neutral acts on how investigators' outcomes are used in retraining.
Under All Governance Targets, the same five meet the targets with Store less data at its stronger setting, for 13 units. At that level, four tools work at reduced strength unless Understand the system is also on. They are Vet connections, Review the riskiest first, Store less data, and Peer sharing rules.
Of the 100,303 combinations within the budget, 169 meet the targets under Service and Safety Targets, and 63 under All Governance Targets.
More is not better here. Every tool at its strongest setting at once closes every failure pathway, but costs 41 units, far over the budget. It also leaves the ranked list no longer clearly helping the investigators' work, so it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Rotterdam-class fraud risk-scoring model network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This network follows the welfare-fraud risk model documented in the Rotterdam case file. It does not rebuild the actual model or its inputs.
- assumed
The network assumes two things about how mistakes pass between people and within the model. Investigators share case habits, and one model's skew repeats across every score. It also draws a check where investigators challenge the list. The sources show no such challenge in use. The scrutiny they record came from the city's audit office and from journalists.
- baseline
The network assumes a retraining loop: past dossiers and the outcomes investigators recorded were used to train the next version of the model. The case file says journalists obtained the model and its training data. Lighthouse Reports' published method says the city retrained the model each year with new data.
- assumed
The network draws the model's inputs as a part of their own. The Suspicion Machines investigation documented about 315 inputs, including age, gender, language, and neighbourhood. Mistakes in those inputs count in the network when they pass to the model. Reviews of risk-scoring systems like this one criticise scoring people from records collected for other purposes. The harm to particular groups is recorded in the case file, outside the network.
- assumed
The ranked risk list is drawn on the pathway from the model to the investigators. The case file describes a model that scored recipients for investigation priority. In the Lab the list is a label on that pathway. The pathway from the model to the investigators carries its effect.
- assumed
The case file records whom this model's scores were skewed against. The network traces how mistakes pass between the model, the investigators, and the dossiers, not between groups of people. It estimates no unequal harm to the people served.
What this example does not show
Show all 1 limitation
- The documented harm is a disparity in whom the model flagged: its scores skewed against women, parents, and people with limited Dutch. The Lab traces how mistakes pass between the model, the investigators, and the dossiers, not between groups of people. It estimates no unequal harm to the people served. That disparity is documented in the case file, and it is measured outside any network like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.
empirical- Government Rekenkamer Rotterdam, Gekleurde technologie: onderzoek naar het gebruik van algoritmes door de gemeente Rotterdam (2021) https://www.rekenkamers.nl/rapport/gekleurde-technologie/
- Investigative Lighthouse Reports, Suspicion Machines (2023) https://www.lighthousereports.com/investigation/suspicion-machines/
- Investigative WIRED / Lighthouse Reports, Inside the suspicion machine (2023) https://www.wired.com/story/welfare-state-algorithms/
- Investigative Follow the Money, How a fraud algorithm learned to suspect vulnerable groups https://www.ftm.eu/articles/algorithm-rotterdam-dissected
- Advocacy Racism and Technology Center, Rotterdam welfare fraud algorithm was biased https://racismandtechnology.center/2023/03/17/racist-technology-in-action-rotterdams-welfare-fraud-prediction-algorithm-was-biased/
Documented risk-scoring deployments computed scores from multi-agency administrative records originally collected for other purposes, which is the data-protection critique recorded in independent reviews of these systems.
empirical- Government evaluation Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf
- Investigative Eubanks, Automating Inequality (2018); The Nation, Want to Cut Welfare? There's an App for That https://www.thenation.com/article/archive/want-cut-welfare-theres-app/
- Investigative Lighthouse Reports, Suspicion Machines (2023) https://www.lighthousereports.com/investigation/suspicion-machines/
Where this connects
Institutional pressures in this domain
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Public benefits & eligibility domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Mark AI-written records — Provenance labeling
- Keep skills sharp — Deskilling-arrest mandate
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Vet connections — Connection authorization
- Review the riskiest first — Risk-tiered oversight
- Upgrade model — Improve the model
- Keep prompts neutral — Framing and mirroring reduction
- Store less data — Data minimization
- Peer sharing rules — Peer-edge governance
- Assign a challenger — Structured dissent
- Escalate checks — State-feedback vigilance
Documented case histories
- Rotterdam welfare-fraud risk model
- Michigan MiDAS
- Robodebt (Australia)
- Indiana / IBM eligibility modernization
- Arkansas ARChoices / ARIA
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- SyRI (Netherlands)
- CNAF benefit-fraud risk score (France)
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Udbetaling Danmark data-driven control (Denmark)
- BOSCO (Spain)
- Serbia Social Card (Socijalna karta)
- UK DWP Universal Credit Advances fraud model
- ID.me identity verification as an unemployment eligibility gate
- Medicaid unwinding: automated ex parte renewal at population scale
- INSS auto-analysis: when the productivity metric makes denial the fastest way out
- Samagra Vedika
- Workforce Australia Targeted Compliance Framework: automated payment sanctioning after Robodebt
- NYC MyCity business chatbot
- Nevada DETR generative-AI unemployment appeals
- Tennessee TennCare TEDS