PAN Lab example
Amsterdam Smart Check
The governed exit: a fair welfare screener that shipped every safeguard but one
Slimme Check flagged welfare applications for investigation. Its bias, reduced on past data, shifted to other groups in a live pilot. Amsterdam halted it.
See more
Slimme Check (Smart Check) was a machine-learning model the City of Amsterdam built to screen new applications for social assistance. It scored each application and labeled some investigation-worthy, meaning worth investigating for possible fraud. The label was advice to the enforcement staff who reviewed it, never a decision.
How it worked
The city began designing Slimme Check in 2019, and the effort ran about five years. It was an Explainable Boosting Machine, a kind of machine-learning model built so people can see how each input moves the score.
It scored each application on 15 behavioral and administrative inputs. They came from the city's personal-records database, with citizen ID, address, and housing situation, and from Income Services systems, with prior applications, assets, and income. The city deliberately left out gender, nationality, and age, and inputs known to stand in for them, such as postal code. Slimme Check learned from about 3,400 past enforcement investigations.
A score above a set cut-off produced an advisory investigation-worthy label. Three people in separate roles then took over. An enforcement employee reviewed each label. A separate investigator carried out any detailed investigation, with powers that include access to bank records and home visits. A different Income Services official made the final decision on the benefit.
Staff were trained, and given written instructions, not to rely on the label too much. The Dutch national algorithm register calls this meaningful human intervention, not automated decision-making.
The safeguards before the pilot
The case file calls these safeguards as complete as the period offered. The city carried out a data-protection impact assessment, and an IAMA, a Dutch assessment of an algorithm's effects on human rights. It also carried out an AI impact assessment and a quick security scan against the Dutch government's information-security baseline. The city's data protection officer and its Personal Data Commission reviewed the design.
The city commissioned a review from Deloitte, about 35,000 euros by the city's own estimate, and consulted the University of Amsterdam's Civic AI Lab. It published entries in both the Amsterdam and the Dutch national algorithm registers.
It also convened a citizen Participation Council of up to 15 benefit recipients and interest-group representatives. The council was briefed in depth in late 2021. MIT Technology Review dates the first briefing to the fall of 2021, and the city's research platform says the end of 2021.
In March 2022, the council concluded the project should be stopped. It cited variables such as year of birth that could discriminate, disproportionate data processing, and use of an algorithm without transparency or democratic control. The city removed some variables and strengthened data safeguards. It overrode the core advice to stop and went ahead with a pilot.
The bias test and the fix
A wrongly flagged application is one labeled investigation-worthy when there was nothing to find. Internal testing in May 2022 found the model wrongly flagged non-Dutch applicants about 30 percent more often. It wrongly flagged non-Western applicants nearly twice as often, and men about 14 percent more often.
The team then reweighted the training data, changing how much each past case counted. On past data, this made wrongful flags roughly equal between Western and non-Western applicants.
What the pilot showed
In 2023, Slimme Check ran in a pilot on nearly 1,600 real applications. The city's research platform gives the window as April to August 2023. The investigation analyzed pilot data from June to August 2023.
The fairness did not carry over. Dutch nationals and women were now more likely to be wrongly flagged. In the investigation's analysis, women were about 22 percent more likely. A new bias against applicants with children appeared, which the city had not publicized. Slimme Check flagged more applications than the caseworker process, not fewer. It was no better than caseworkers at picking out the applications that did warrant investigation. The city's internal claim before the pilot, an accuracy edge of about 20 percent over caseworkers, did not hold.
The city halted the project. The Dutch national algorithm register records the deployment ending in September 2023 and lists Slimme Check as out of use. In late November 2023, the alderman responsible, a member of the city's executive, announced the end publicly. He said he could not have justified continuing a pilot that showed the algorithm contained substantial bias.
Where the figures come from
The group figures and the flag counts are not city publications. Journalists from MIT Technology Review, Lighthouse Reports, and Trouw computed them, with Pulitzer Center support. They worked from confusion matrices the city provided: tables counting correct and wrong flags for each group. The city ran the journalists' code on real applicant data and returned only totals, under an arrangement that complied with the GDPR, the EU's data-protection law.
The journalists had cooperative access to several model versions and the city's code. In June 2025 they published their full method, both model files, the bias-evaluation code, and the city's tables on GitHub. The model files are the original and the reweighted one.
Which groups were wrongly flagged more often reversed between the past data and the live pilot.
What each side says
The caseworker process Slimme Check was meant to improve on was itself flawed. Caseworkers were more likely to wrongly flag Dutch nationals and women. More than half of the investigations caseworkers started found no wrongdoing. Fraud is estimated at roughly 3 percent of applications. The city had about 35,000 welfare recipients at the end of 2024.
Some project staff argued for more testing. They said going back to the caseworker process was itself "a decision to go back to" a biased system. Whether the halt was a governance success, or the abandonment of a tool that could be fixed, is a live disagreement.
The Racism and Technology Center, an advocacy group, reads the same events as the city overriding its citizen panel's clear advice. It argues that adjusting the model for responsible AI missed the deeper question: whether such a system should be built at all.
What this case asks
The case file argues that fairness is not a property you install in a model and then own. It is a relationship between a model and the live population it scores. A check before launch can test that relationship only against past data.
Nearly everything here was built before the pilot. The case file argues the missing piece was a standing check of live results, on a schedule, with the power to halt. The pilot's bias evaluation was the city's own, run once after the pilot. Outside investigators measured the reversed bias many months later.
What this network is drawn from
This network follows the pattern the sources describe, modeled on Amsterdam's Slimme Check pilot. It is not a reconstruction of Slimme Check itself. It shows Slimme Check, the pre-pilot assessments and reviews, the Participation Council, the label reviewer, the investigator, the decision officer, and the applicant and investigation records. It also keeps two checks on the map: the council's overridden advice to stop, and a recheck against live applications, which the city did not have.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on. A mistake here is, for example, a wrong investigation-worthy label or a wrong finding entered in the record.
This case has a budget of 10 units. Each tool costs the same at every target level. The pressure Monitoring goes stale is on from the start: no one has to act on what monitoring shows.
The work here is screening applications for investigation. Explore (No Targets) sets no targets. Under Service Targets Only, the targets ask that mistakes stop building on one another, and that Slimme Check keep helping staff get the work done. Before any tool is used, the network sits at a tipping point: mistakes are on the edge of building on one another. A self-correcting network is one where they fade out instead. Assign a challenger, Peer sharing rules, or Mark AI-written records each meets the targets on its own. Many pairs of other tools meet them too.
Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, and to keep the work from falling behind, among other targets. Before any tool is used, seven failure pathways are open.
One carries Slimme Check's label to the label reviewer. Two carry each case from the reviewer to the investigator, and from the investigator to the decision officer. Two carry the records out, to Slimme Check as its inputs and to the label reviewer. Two carry findings and decisions into the records.
Every combination within the budget that meets these targets has four tools. Escalate checks closes the label pathway. Mark AI-written records closes the two pathways out of the records. Peer sharing rules or Assign a challenger closes the two hand-offs between staff. Gate record entries or Store less data closes the two pathways into the records. Each of these four combinations costs 9 units. The last unit can buy the stronger setting of Peer sharing rules, Assign a challenger, or Mark AI-written records.
All Governance Targets asks more of the service, and the same four combinations still meet it. Pause AI on alarms also closes the label pathway, but it leaves the work falling behind, so it meets the targets at no level. Check with a second model adds the check named Recheck against live applications. But staff then rely more on the labels, catch fewer mistakes, and fall behind. No combination at the two higher levels includes it.
More is not better here. Using every tool, each at its strongest setting, costs 45 units, four and a half times the budget. It closes every pathway, but the work falls behind. So it meets the targets at none of the three levels that set them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Amsterdam-Slimme-Check-class governed-exit screening pipeline network: 7 components and 15 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
This network follows the pattern the Amsterdam Slimme Check case file describes. A city screening pilot had nearly every recommended safeguard, and the city itself halted it when bias came back in live use. It does not reconstruct the actual tool.
- assumed
The network assumes mistakes pass along the chain of three staff roles. A label confirmed by mistake leads to an investigation, and its finding shapes the final decision. It assumes the decision officer cross-checks each case. It also assumes one model scoring every application makes a flaw repeat across all of them. That is why three separate people, each reading one case, could not see it.
- assumed
The network draws nearly every recommended safeguard before the pilot as in place. These are the data-protection and human-rights assessments, the outside review, the academic advice, the two register entries, and the citizen council. Reweighting the training data, changing how much each past case counted, made wrongful flags roughly equal between Western and non-Western applicants on past data. The network keeps two checks on the map so you can see a tool add them. One is the council's advice to stop, which the city overrode. The other is a standing recheck against live applications, which the city did not have.
- assumed
Slimme Check reads city records and income data as its inputs. The network also draws it entering its score in the record. That pathway is inferred: the sources describe only the label shown to reviewers, and no score written back. People review, investigate, and decide, and a person enters each final decision in the permanent record. So whether a mistake reaches a decision depends on the check of each case holding. A check of one case at a time cannot see a skew that shows only across all applications.
- assumed
Applicants, and whether they receive social assistance, are not part of this network. The group figures are the investigating journalists' calculations on summary tables the city provided. The city produced those tables under an arrangement that complied with the GDPR, the EU's data-protection law. They are not city publications. Which groups were wrongly flagged more often reversed between the past data and the live pilot. No unequal harm to applicants is estimated here. Those outcomes are recorded in the case file and measured outside this network.
What this example does not show
Show all 3 limitations
- Applicants, and whether they receive social assistance, are not shown here. The Lab shows only how mistakes pass between people, tools, and records inside the institution. Those outcomes are recorded in the case file and measured outside this example.
- Before the training data was reweighted, non-Western applicants were nearly twice as likely to be wrongly flagged, and men about 14 percent more likely. Reweighting changed how much each past case counted. In the live pilot, women were about 22 percent more likely, and Dutch nationals and applicants with children were also wrongly flagged more. These are the investigating journalists' calculations, not city publications. They worked from tables of correct and wrong flags that the city produced, under an arrangement that complied with the GDPR, the EU's data-protection law. Which groups were wrongly flagged more often reversed between the past data and the live pilot. The figures give context for decisions. They are not findings about causes. No unequal harm to applicants is computed here.
- The cost of roughly 535,000 euros and the claim before the pilot of an accuracy edge of about 20 percent over caseworkers are the city's own internal estimates. Neither was audited. The accuracy edge did not hold in the live pilot. The halt is reported as the city's own decision. The national register gives September 2023 as the end date, and the termination was announced publicly in November 2023. People inside the project disagreed about whether stopping was a governance success or gave up a tool that could be fixed. The case file carries that disagreement. This example does not settle it.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The City of Amsterdam spent roughly five years and an estimated EUR 535,000 building a deliberately fair, explainable welfare-fraud screening model with nearly every recommended pre-deployment safeguard in place - a bias audit, training-data reweighting that approximately equalized wrongful-flag rates on retrospective data, a data-protection assessment and a human-rights assessment, external and academic review, a citizen panel, and dual algorithm-register transparency - and discontinued it after a 2023 live pilot on nearly 1,600 applications. In the investigating journalists' analysis of aggregate data the city provided, the group disparities re-emerged inverted on the live pilot, now more likely to wrongly flag Dutch nationals, women, and applicants with children, with the tool flagging more applications than the analog process and no better than caseworkers at finding genuine cases. The Dutch national algorithm register records the deployment ending September 2023 and lists it out of use, and the responsible alderman announced the halt in November 2023.
empirical- Investigative Braun, Geiger, Amsterdam Fair Welfare AI (Inside Amsterdam's high-stakes experiment to create fair welfare AI) (MIT Technology Review, with Lighthouse Reports and Trouw, 2025) https://www.technologyreview.com/2025/06/11/1118233/amsterdam-fair-welfare-ai-discriminatory-algorithms-failure/
- Investigative Lighthouse Reports, Amsterdam's 'Smart Check' welfare-fraud model: fairness methodology (with Trouw and MIT Technology Review, supported by the Pulitzer Center) https://www.lighthousereports.com/methodology/amsterdam-fairness/
- Government Algoritmeregister (Dutch national algorithm register), Onderzoekswaardigheid: Slimme check levensonderhoud (Gemeente Amsterdam) (2023, last modified 2025) https://algoritmes.overheid.nl/nl/algoritme/gm0363/95794697/onderzoekswaardigheid-slimme-check-levensonderhoud
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Escalate checks — State-feedback vigilance
- Review on schedule — Oversight cadence & retrospectives
- Keep skills sharp — Deskilling-arrest mandate
- Assign a challenger — Structured dissent
- Peer sharing rules — Peer-edge governance
- Check with a second model — Cross-model verification
- Vet connections — Connection authorization
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Mark AI-written records — Provenance labeling
- Pause AI on alarms — Deployment circuit-breaker
Documented case histories
- Amsterdam Smart Check
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot