PAN Lab example
New Zealand MSD Predictive Risk Modelling
Halted before it ran: a national child-risk model
New Zealand's Ministry of Social Development commissioned a model to score newborns' maltreatment risk. A minister halted its study, and the tool never went live.
See more
The predictive risk modelling tool was a statistical model commissioned by New Zealand's Ministry of Social Development. It would have scored a child, at or shortly after birth, for the likelihood of a substantiated maltreatment finding by age five. A substantiated finding is a decision by the child protection agency that maltreatment occurred.
What is decided
The tool was designed to give frontline social workers a score for each child, to weigh alongside their own assessment. Tim Dare's 2013 ethical review for the ministry records that the score was to supplement a full clinical assessment, not replace it.
The model computes the score from linked records. They are income-support benefit records and child protection records from Child, Youth and Family, the child protection service. Which children get scored is a policy choice, not the social worker's.
How the model was built and tested
The development work was led from the University of Auckland, with co-authors at Auckland University of Technology and the University of Southern California. It was published in 2013 in the American Journal of Preventive Medicine.
The developers used a 2012 sample of 57,986 children. They selected 132 inputs from 224 they had built. The area under the curve, a standard measure of how well a score ranks cases, was 76 percent. Among the tenth of children the model ranked highest, 47.8 percent had a substantiated finding by age five.
Those figures come from the development data. They were never tested in the field, because the study that would have tested them was the one the minister stopped.
The reviews before any use
The design went through an unusual set of reviews before any use. Tim Dare wrote an independent ethical review in 2013. He judged that the remaining concerns "may plausibly be regarded as outweighed by the very considerable potential benefits".
A dedicated Maori ethical review was also done. A Privacy Impact Assessment, a written review of a design's privacy risks, was discussed with the Office of the Privacy Commissioner. The sources do not say what the Maori review or the privacy assessment concluded.
The study that was stopped
In November 2014 the ministry lodged an ethics committee application for a two-year study. It would have scored about 60,000 newborns by linking maternity, benefit, and child protection records. Then it would have watched, without acting, whether the high-risk predictions came true.
The incoming Social Development Minister, Anne Tolley, withdrew the application. She wrote on the briefing papers: "Not on my watch! These are children not lab rats." The Children's Commissioner took the same view: "no ethics committee in New Zealand is going to wear that".
What happened next
Testing was narrowed to about 20 historic cases, with identifying details removed. The tool was never used in live casework. As of 2026, frontline workers at Oranga Tamariki, which the sources name as the later home of this work, still do not use predictive modelling.
What critics raised
The model's target, a substantiated finding, is itself an agency decision, not a measure of harm. Past findings were also among its inputs. So the patterns in earlier decisions would return in every future score.
The records were collected for other purposes, and the families did not agree to their use for scoring. Reviewers and critics also raised the concern that both the outcome and the inputs could carry existing bias against Maori and benefit-receiving families.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A pathway is closed when mistakes stop passing along it. The work along it goes on.
This case has a budget of 9 units. Explore (No Targets) sets no targets. Under Service Targets Only, the targets are already met before you use any tool. That level asks that mistakes not build on one another, and that the model stay useful to the work.
Under Service and Safety Targets and All Governance Targets, the targets can also be met. Both levels ask you to close every failure pathway and refill the Privacy gauge, among other targets. Four pathways start open: the score given to social workers, workers recording decisions, the model reading the linked records, and workers reading a child's history.
The cheapest way costs 7 units: Escalate checks at 2, Store less data at 3, and Mark AI-written records at 2. Escalate checks is the one tool offered here that closes the pathway where the score is given to social workers.
One other way leaves out Mark AI-written records. It uses Escalate checks, Vet connections, and Store less data with Understand the system, for all 9 units. At these two levels Understand the system costs 4, and it takes 1 unit off each of the other three.
Counting stronger settings, eight distinct sets of tools meet the targets at each of these two levels. Every one includes Escalate checks and Store less data.
Require sign-off, the control that decided the real case, is not needed at any level. The network already includes the review step that held.
More is not better here. Every tool at its strongest setting at once costs 35 units, far over the budget. It also leaves the model no longer clearly helping the work, so it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the NZ-PRM-class halted-before-deployment risk model network: 4 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
This network follows the pattern in the case file on New Zealand's predictive risk modelling tool, which was stopped before use. It does not rebuild the actual tool.
- baseline
The tool was never used in live casework. Its published accuracy figures come from development data and were never tested in the field. This network shows the shape the design would have taken, not a system that ran.
- assumed
The network assumes that past decisions return in future scores. The model's target, a substantiated maltreatment finding, is itself an agency decision. Past substantiation findings were also among its inputs. Reviewers and critics raised this concern.
- assumed
The ethical and privacy review part stands for the reviews done before any use. They were an independent ethical review, a dedicated Maori ethical review, and a Privacy Impact Assessment discussed with the Privacy Commissioner. They checked the design before use, not each child's case.
- assumed
Social workers pick up screening habits from one another. One national model would repeat its blind spots across every child. Second reads by colleagues and supervisors are a check among workers. No independent second model was ever built to check the score.
- baseline
The network treats the review step between the model and social workers' casework as a check that held. The sources show the reviews ran, and an accountable minister refused to authorize the study. The tool never scored a live case.
- assumed
Reviewers and critics raised the concern that both the outcome and the inputs could carry existing bias against Maori and benefit-receiving families. The case file documents that concern. This Lab follows how mistakes pass between the parts of an organization, not whom they fall on. It estimates no unequal harm to the families served.
What this example does not show
Show all 2 limitations
- This tool was never used in live casework. The network shows the shape its design would have taken, not a system that ran. Its often cited accuracy figures come from development data and were never tested in the field.
- Reviewers and critics raised a concern about bias: the outcome and the inputs could carry existing bias against Maori and benefit-receiving families. The case file documents that concern. This example follows how mistakes pass between the parts of an organization, not whom they fall on. It estimates no unequal harm to the families served, and any such harm would have to be measured outside the Lab.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
New Zealand's Ministry of Social Development commissioned a child-maltreatment risk-modelling tool that, on a 2012 development sample of 57,986 children and 132 selected variables, reported an area under the ROC curve of 76% and a top risk decile in which 47.8% had a substantiated maltreatment finding by age five; those figures come from development data rather than field performance, the tool was never operationally deployed, and a proposed two-year study that would have scored about 60,000 newborns was halted by the incoming Social Development Minister, who annotated the briefing papers 'Not on my watch! These are children not lab rats.'
empirical- Academic Vaithianathan, Maloney, Putnam-Hornstein, Jiang, Children in the Public Benefit System at Risk of Maltreatment: Identification Via Predictive Modeling (American Journal of Preventive Medicine, 2013) https://csda.aut.ac.nz/__data/assets/pdf_file/0019/11926/children-in-the-public-benefit-system-at-risk-of-maltreatment1.pdf
- Investigative NZ Herald, Anne Tolley scraps 'lab rat' study on children (2015) https://www.nzherald.co.nz/nz/anne-tolley-scraps-lab-rat-study-on-children/C7GIGYW2467HG327FKXFRJDPEM/
- Investigative Otago Daily Times, Call to stop child abuse risk modelling study (2015) https://www.odt.co.nz/news/national/call-stop-child-abuse-risk-modelling-study
- Investigative Mordaunt, Child protection workers are under pressure in NZ. Can predictive modelling help? (The Conversation, 2026) https://theconversation.com/child-protection-workers-are-under-pressure-in-nz-can-predictive-modelling-help-278298
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Escalate checks — State-feedback vigilance
- Require sign-off — Conformity assessment gate
- Review on schedule — Oversight cadence & retrospectives
- Vet connections — Connection authorization
- Store less data — Data minimization
- Mark AI-written records — Provenance labeling
- Keep skills sharp — Deskilling-arrest mandate
- Assign a challenger — Structured dissent
- Review the riskiest first — Risk-tiered oversight
- Upgrade model — Improve the model
- Understand the system — Understand the system
Documented case histories
- New Zealand MSD Predictive Risk Modelling
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- Gladsaxe model