Skip to content

PAN Lab example

ProKid (Netherlands)

The colour and the record: a child risk-profiler

ProKid, a Dutch police instrument, sorted children under 12 into colour risk bands from police records, including children recorded only as victims or witnesses.

See more

ProKid Signaleringsinstrument 12-, or ProKid, was a rule-based instrument built by the Gelderland-Midden police. It sorted children under 12 who appeared in police records into four colour risk bands: white, yellow, orange, and red. It computed each colour from up to twelve years of police records. The sources do not say exactly what the colours were meant to predict.

How it was used

Police quality controllers, the police staff who decided what to do with each flag, received the colour flags. They decided which children to refer to the Youth Care Agency, Bureau Jeugdzorg. The agency assessed whether help was needed and arranged it. In Amsterdam, children with a red flag were referred straight to the agency.

The instrument counted incidents in which the child was a suspect, a witness, or a victim. It also counted reports at the home address and about the people living there. So a child recorded only as a victim could be given a higher colour. Reports at the home address could raise a child's colour too.

The controllers re-checked the criteria by hand before referring anyone. The sources do not say what they compared the criteria with.

What the evaluation found

The government commissioned an evaluation of the pilots. It ran from 2009 to 2011 in four police regions.

In one three-month window, 902 of 2,444 red, orange, and yellow flags were system or registration errors, or based on irrelevant incidents. That is 36 percent, and 53 percent in Amsterdam-Amstelland.

The evaluation concluded that none of the four regions had a well-functioning instrument. The colour bands carried little weight in practice, and they showed no difference in the intensity of help children needed. Outside Amsterdam, all three flagged bands were handled the same way. Decisions turned on the last incident and on the family's protective factors, meaning what in the family helps keep a child safe.

What reporting asked

Reporting by Sargasso in 2012 asked whether a wrong entry stays in the records, uncorrected, and counts toward the next child's score. The pilot evaluation could not settle that question. The sources do not say whether any wrong entry was corrected.

The children's-rights critique

A children's-rights critique objects to profiling children at all. It objects to profiling victims as possible offenders, and to treating a child's whole recorded life as evidence. In its view, a tool can predict future police contact well and still do both.

What came after

A later successor, ProKid Plus, was an automated model built on police data. Its validation reported an area under the curve of about 0.83 for predicting future violent or property offending. That measure shows how well a score ranks a real case above a non-case, where 0.5 is a coin flip.

In answers to parliamentary questions in December 2022, the Minister of Justice and Security reported that ProKid Plus had been used once. That was as a pilot input to a programme in Amsterdam called Top400, in July 2016, which the sources read here do not describe. It was no longer used, and no further successor had been built.

ProKid ended not in a court ruling but in disuse.

Who oversaw it

There was no dedicated oversight of police algorithms when ProKid was deployed. By 2022, the Justice and Security Inspectorate, the Data Protection Authority, and the Court of Audit were named as overseers of police algorithms.

Why this case matters

The case file reads ProKid as a case about how a child's record is built. The colour carried little weight in practice, so the score was never the main thing. The risk sat earlier, in the records the colour was computed from: which children were profiled, from which records, and kept for how long. It also sat in whether a wrong record was ever corrected before it counted again.

The case file names the question accuracy measures never answer. It is not whether the score is right, but whether the record should have been built at all.

Where the facts come from

The network draws on seven sources. Two are the DSP-groep pilot evaluation for the WODC (2011) and Dimitri Tokmetzis's reporting for Sargasso (2012). Two are Karolina La Fors's study of children's rights (2015) and Praktikon's study of ProKid Plus (2016). The others are a 2017 validation study of the ProKid tool, the minister's 2022 answers to parliament, and Fair Trials' 2021 report, Automating Injustice.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it may go on.

This case has a budget of 11 units. Before any tool is used, four failure pathways are open: Police history scored in, Decisions written to records, Police records as context, and Flagged children referred on. The mistakes are not contained, meaning they build on each other rather than being corrected.

The Lab starts with its Dynamics setting at Both, where side-effects and lingering effects are both on. The counts below use that setting unless they say otherwise. Service and Safety Targets fixes Dynamics at Side-effects, and All Governance Targets at Both. What follows for those levels holds at those settings.

Explore (No Targets) sets no targets. There, one tool is enough to keep the mistakes contained. The cheapest are Mark AI-written records and Peer sharing rules, at 2 units each.

Under Service Targets Only, the targets can be met. This level also asks for the instrument to be helping the work. Mark AI-written records does it alone, for 2 units. Upgrade model at its stronger setting also does it alone, for 5 units. With Dynamics at Off, Vet connections alone does it for 3 units, and Upgrade model does not. Of the 345 sets of tools that fit the budget, 131 meet these targets.

Understand the system pays for ongoing study of what the deployment is really doing. It costs 3 units under Explore and Service Targets Only, and 4 under the other two levels. While it is on at its standard setting, Mark AI-written records, Vet connections, Peer sharing rules, and Upgrade model each cost 1 unit less. At its stronger setting, which costs 6 units, each of the four costs 1 unit at its standard setting.

Under Service and Safety Targets, the targets are not fully addressable with the available tools. This level also asks you to close every failure pathway, and to keep the Work getting done gauge from reading strained. Closing all four takes Peer sharing rules, for Flagged children referred on, and Store less data, for Decisions written to records.

Every combination that closes all four and contains the mistakes falls short in one of two ways. Without Understand the system, the work reads strained. With it, the instrument no longer helps the work enough, and in some combinations the work still reads strained.

Under All Governance Targets, the targets are not fully addressable either, for the same reason. That level asks for the instrument to be clearly helping the work.

Money is not what stands in the way. With the budget set aside, no combination meets the targets at either of those two levels.

More is not better here. Using every tool at its strongest setting costs 28 units. It contains the mistakes, closes every failure pathway, and keeps up with the work. But the instrument then no longer helps the work, so it meets the targets only under Explore (No Targets).

Stylized model of a documented deploymentChild welfare & family services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the ProKid-class child risk-profiling instrument network: 4 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    This network follows the child risk-profiling pattern documented in the ProKid case file. It does not rebuild the actual instrument.

  • baseline

    The network places the main risk in the police records the colour was computed from, not in how controllers acted on it. The instrument drew on up to twelve years of police records. The pilot evaluation found 902 of 2,444 red, orange, and yellow flags in one three-month window were system or registration errors, or based on irrelevant incidents. Reporting asked whether such entries stay in the records and count toward the next score. The Allegheny Family Screening Tool, a Pennsylvania county's score for child-maltreatment calls, differs. Its case file finds that screeners' overrides narrowed the racial gap the score alone would have produced.

  • assumed

    The network assumes the colour counted for little in the controllers' decisions. The pilot evaluation found the colour bands carried little weight in practice. Outside Amsterdam, all three flagged bands were handled the same way. Decisions turned on the last incident and on the family's protective factors.

  • assumed

    The network draws checks between people and between records. Quality controllers re-checked the criteria before referring anyone, and the network assumes they also shared screening habits. The sources describe no step that corrected a wrong entry before the next score. Reporting asked about exactly that step.

  • assumed

    The documented harm falls on the children the instrument flagged, including those flagged in error and children recorded only as victims or witnesses. A children's-rights critique objects to profiling them at all. The sources document no ethnic-bias finding specific to ProKid, and none is claimed here. The network traces how mistakes pass between the instrument, the people, and the records, not outcomes for children. It estimates no unequal harm to the people served. The case file records that harm, outside the network.

What this example does not show

Show all 1 limitation
  • This example shows how mistakes pass between the instrument, the police, the Youth Care Agency, and the records. It does not show what happened to children. The documented harm falls on the children the instrument flagged, including those flagged in error and children recorded only as victims or witnesses. A children's-rights critique objects to profiling them at all. The sources document no ethnic-bias finding specific to ProKid, and none is claimed. This example estimates no unequal harm to the people served. The case file records that harm, outside this example.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange, and yellow child-risk flags (902 of 2,444 over three months across four police regions, rising to 53% in Amsterdam-Amstelland) were system or registration errors or based on irrelevant incidents, and that in none of the four regions was there a well-functioning instrument.

    empirical
    • Government evaluation DSP-groep for the WODC (Abraham, Buysse, Loef & van Dijk), Pilots ProKid Signaleringsinstrument 12- geevalueerd (2011) https://repository.wodc.nl/handle/20.500.12832/1832
    • Investigative Dimitri Tokmetzis (Sargasso), Hoe de politie duizenden risicokinderen produceert (2012) https://sargasso.nl/hoe-de-politie-duizenden-risicokinderen-produceert/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Child welfare & family services domain page.

Levers available here and the patterns behind them

Documented case histories