PAN Lab example
What Works for Children's Social Care ML pilots
The bar it never cleared: a child-welfare prediction pilot
A research centre's 32 models predicted which children's cases would escalate. None was right often enough to clear its published bar, so none went live.
See more
What Works for Children's Social Care, a research centre funded by England's Department for Education, built 32 machine-learning models in-house with four local authorities. These are the councils that run children's services. Each model used past case records to predict one kind of escalation in a child's case, such as a child protection plan. None was used in live casework.
How the project ran
The project took about 18 months. The programme appears to have recruited up to five local authorities, with one withdrawing. The team built the models from three to seven years of the authorities' past case records.
It tried three kinds of model. Decision trees ask a series of yes-or-no questions about a case. Logistic regression adds up weighted factors. Gradient boosting combines many small decision trees.
Some builds also read free-text case notes using natural-language techniques, which let a computer read written text. The notes were pseudonymised, meaning names and other identifiers were replaced.
What the models predicted
The eight outcomes included a referral to statutory services within 12 months of early help. Early help is support offered to a family before statutory social care, the care a local authority must provide by law, becomes involved.
Other outcomes were escalation to a child protection plan, and a child becoming looked after, meaning taken into the local authority's care.
The bar and the result
Before the work began, the team published a success bar. Each model had to reach an average precision of 65%. Average precision is a standard score of how many of a model's flags are right, averaged across its possible cut-offs. A model flags a child's case when its risk score for that case passes a cut-off, a level the team chooses.
None of the 32 models cleared it. The best single model reached only about 42%. At one chosen cut-off, it missed about 79% of the children whose cases actually escalated.
Taken together, the models missed around four in five children who were genuinely at risk. When they flagged a child, they were wrong roughly six times in ten.
The team reported this in September 2020, in "Machine learning in children's services: does it work?" The models were never placed in front of social workers in live casework.
Why the models fell short
The authors gave four reasons. The outcomes were rare. Siblings' cases were similar to one another. There was little past data to learn from. And care pathways, the routes children take through services, were complex.
The case file draws a lesson from this. When serious escalation is rare, accuracy has a ceiling that no change of model climbs. The honest move is a gate that refuses deployment, not another model.
What social workers thought
A survey of 129 social workers, carried out for the project, found low support. About 26% supported using predictive analytics to pick out families for early help. About 34% thought it should not be used at all.
What happened next
What Works for Children's Social Care did not deploy the models anyway. It published the negative result. It also pressed the wider sector to disclose how effective the tools already in operational use are. Its executive director, Michael Sanders, concluded it was time for the sector to "stop and reflect."
A companion ethics review by the Alan Turing Institute and the Rees Centre, from January 2020, recommended national standards. It found systems that devalue person-centred approaches "not ethically permissible."
The ethics review also flagged a concern about models like these. Models trained on past intervention decisions learn recorded practice, not underlying risk.
Why this case stands out
Other systems in these case files were stopped before use by other means, such as a minister's refusal or a missing legal basis. This one was stopped by a rule written down in advance: a public effectiveness bar the models had to clear to go live. None did, so none was put in front of a social worker. It is the rare case where the question "does it actually work?" was asked first, in public, and answered honestly.
The case also shows transparency as a control, meaning a safeguard that keeps a system in check. Publishing a negative result, and pressing tools already in use to disclose their effectiveness, is what makes an evidence bar enforceable across a market.
What this example shows
The network shows the pilots as if the models had been put to their intended use, with scores offered to social workers at intervention decisions. In the pilots, the scores stayed inside the research evaluation. The example rehearses whether to deploy them, a question the evaluation answered no.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Here a mistake is, for example, a wrong prediction about a child's case. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 11 units. Understand the system costs 3 units under Explore and Service Targets Only, and 4 under the two higher levels. Its stronger setting costs 6 units. Every other tool costs the same at every level.
While Understand the system is on, Store less data, Upgrade model, Escalate checks, and Pause AI on alarms each cost 1 unit less. At its stronger setting the cut is 2 units, but no tool drops below 1 unit.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets are already met before any tool is used. Mistakes are caught faster than they are copied, and the models help the work. More than 580 different sets of tools within the budget meet them, including using none.
Under Service and Safety Targets and All Governance Targets, the targets can be met. Both levels ask you to close every failure pathway, among other targets. Before any tool is used, three are open: Predictions scored against the bar, Case history trains the models, and Social workers' decisions recorded.
Closing all three takes three tools. Escalate checks, which raises checking when monitoring flags trouble, closes Predictions scored against the bar. Mark AI-written records closes Case history trains the models. Store less data, which writes and keeps fewer records, closes Social workers' decisions recorded. Together they cost 7 of the 11 units.
In the Lab, Mark AI-written records marks machine-made content in the records. No scores entered the case history in the pilots. So here it stands for marking which recorded decisions a score shaped.
Every set of tools that meets the targets at these two levels includes all three. There are 21 such sets under Service and Safety Targets, and 18 under All Governance Targets. That level also asks the models to be clearly helping, which rules out three of the 21.
Pause AI on alarms also closes Predictions scored against the bar. But it lowers the benefit the models add below what the targets ask, and the work falls behind. So no set that meets them includes it.
With lingering effects on, Store less data and Review the riskiest first work at reduced strength unless Understand the system is also on. Lingering effects is a Dynamics setting in which damage outlasts its cause. It is on by default, and always on under All Governance Targets. Store less data still closes its pathway at reduced strength.
Both higher levels also ask for a full Privacy gauge. Case history trains the models is the one pathway that drains it, because it reuses children's and families' case records. Closing it stops the drain.
Using every tool on offer, each at its strongest setting, costs 36 units, more than three times the budget. It closes every failure pathway. But together the tools lower the benefit the models add below what every level asks. So it would meet the targets at no level.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the WWCSC-class pre-deployment prediction pilot network: 4 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
This example follows the What Works for Children's Social Care machine-learning pilots, as their case file describes them. It does not rebuild the actual research programme or its models.
- assumed
The example shows the models as if put to their intended use, with scores offered to social workers at intervention decisions. In fact the research build never entered live casework. It was meant to support decisions where social workers use wide professional judgment. A survey for the project found low support for such tools among social workers. So the example assumes social workers would often question the scores rather than defer to them. It rehearses whether to deploy them, a question the evaluation answered no.
- baseline
The research team's evaluation is its own part of the network, with a pathway from the models into it. The team scored the models against a published success bar before any decision to go live. Whether the models cleared it comes from the case file. The Lab does not compute it.
- assumed
The network includes the models learning from the case history from the start. The models learned from past intervention decisions, so they learn recorded practice, not underlying risk. The companion ethics review flagged this concern. It is a documented risk in the data, not a measured difference between groups in these models.
- assumed
The sources give no figures on differences in harm between groups of children or families. Any such harm here is a documented risk in how the models learn, not a measured outcome. This example shows only how mistakes pass between the models, social workers, and records. Anything known about harm to children and families is in the case file, and would be measured outside this network.
What this example does not show
Show all 1 limitation
- The models never ran in live casework. This example rehearses whether to deploy them, a question the evaluation answered no. Bias appears here as a documented risk in how the models learn from past decisions, not as a measured difference between groups. The network shows only how mistakes pass between the models, social workers, and records. Anything known about harm to children and families is in the case file, and would be measured outside this network.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.
empirical- Academic Clayton and Sanders, Can Machine Learning Save Children at Risk? (Significance, Royal Statistical Society) (2022) https://academic.oup.com/jrssig/article/19/6/22/7072840
- Trade press Community Care (Turner), 'No evidence' machine learning works well in children's social care, study finds (2020) https://www.communitycare.co.uk/2020/09/10/evidence-machine-learning-works-well-childrens-social-care-study-finds/
- Government evaluation ChildHub (Terre des hommes), Machine learning in children's services: does it work? (library record) (2020) https://childhub.org/en/child-protection-online-library/machine-learning-childrens-services-does-it-work
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Escalate checks — State-feedback vigilance
- Require sign-off — Conformity assessment gate
- Review on schedule — Oversight cadence & retrospectives
- Gate vendor updates — Vendor quality gate
- Pause AI on alarms — Deployment circuit-breaker
- Upgrade model — Improve the model
- Review the riskiest first — Risk-tiered oversight
- Mark AI-written records — Provenance labeling
- Store less data — Data minimization
- Assign a challenger — Structured dissent
- Understand the system — Understand the system
Documented case histories
- What Works for Children's Social Care ML pilots
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model