PAN Lab example
Illinois Rapid Safety Feedback
The alarm that cried wolf: a child-safety scorer
Illinois's child-welfare agency scored children in reports of suspected abuse or neglect. Its tool flagged thousands as extreme risks and missed children who died.
See more
Rapid Safety Feedback was a predictive tool from the nonprofit Eckerd Connects and its for-profit partner, MindShare Technology. Illinois's Department of Children and Family Services, DCFS, used it until December 2017. Each night it scored every child named in a hotline report of suspected abuse or neglect for the chance of death or serious injury within two years.
How it was meant to work
Each score ran from 1 to 100. A high score was designed to prompt a supervisor's review, a "second set of eyes", and coaching for the caseworker.
Eckerd says front-line caseworkers should never get the raw scores, let alone decide on them. It says DCFS supervisors, trained and coached by Eckerd, should review the scores and decide which cases need immediate attention.
What went wrong
DCFS's internal tracking data, released under Illinois public-records law, showed more than 4,100 children scored at 90 percent or higher. That included 369 children under age 9 given a score of 100 percent. It was far more than any caseload could act on. Reporting found caseworkers alarmed and overwhelmed by the alerts.
Meanwhile, children who died in cases DCFS already knew were not flagged as top risk. Semaj Crosby, 17 months old, was found dead after at least ten DCFS investigations. Itachi Boyle, 22 months old, also died without a high score.
The records behind the score
The scores were recomputed each night from DCFS's records, and those records were flawed. They had entry errors. They often failed to link a child's history to siblings or other adults in the home. State law required DCFS to erase investigations closed as "unfounded".
So the tool worked from a patchy, incomplete picture. It scored thousands of children as extreme risks while missing children with long histories with DCFS.
Why the alarm failed both ways
A score is calibrated when it means what it says: a 30 percent score should match about 30 cases in 100. The case file says this failure ran deeper than calibration, because the records behind the score were patchy.
A flag that fires for thousands cannot change where scarce attention goes. It wears people down with alarms. It also makes an unflagged case look cleared. Trusting such an alarm brings false panic and false calm together.
How it came in and how it ended
George Sheldon became DCFS director in 2015, after a run of child deaths. The roughly $366,000 program was central to the reforms he promised. DCFS brought it in through a no-bid arrangement that the state classified as a grant.
In July 2017, a joint report by the Illinois Executive Inspector General and the DCFS Inspector General found that classification was mismanagement. It had sidestepped the state's rules for transparent bidding.
In December 2017 the new director, Beverly "B.J." Walker, ended the predictive scoring. She said it "wasn't predicting any of the bad cases." DCFS kept a smaller Eckerd case-review training program. Walker said 15 staff and three supervisors were using it.
Who caught it
The decisive corrections came from outside DCFS. They were investigative reporting, a public-records release that put the extreme scores in plain view, and the inspectors general's finding on how the program was bought.
Beyond Illinois
Reporting and a later review by the American Civil Liberties Union placed Illinois among several jurisdictions that took up the tool and then dropped it. Alaska, Louisiana, Ohio, and Oklahoma were among them. The scored families were typically never told they were scored.
What this network is drawn from
This network follows the pattern the case file describes. It is not a reconstruction of the actual tool. It shows the scorer, the hotline reports, the list of highest scores, the DCFS caseworkers and supervisors, and DCFS's case records.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.
This case has a budget of 14 units. Before any tool is used, mistakes build on one another across the network, and eight failure pathways are open.
Explore (No Targets) sets no targets. There, and under Service Targets Only, one tool is enough to stop mistakes building on one another: Mark AI-written records, at 2 units. Under Service Targets Only it also meets the service target, which asks that the tool be helping the work. Many other combinations do both.
Service and Safety Targets also asks you to close every failure pathway, among other targets. The targets can be met within the budget. The cheapest way costs 12 units. It uses five tools: Escalate checks, Mark AI-written records, Keep prompts neutral, Store less data, and Check with a second model.
Every way to meet these targets includes all five. Counting each tool's stronger setting as a separate choice, there are 25 ways within the budget. They use 7 different sets of tools.
Understand the system costs 4 units at the two upper levels and 3 units at the two lower ones. Its stronger setting costs 6. While it is on, Mark AI-written records, Store less data, Check with a second model, and Require sign-off each cost 1 unit less. At its stronger setting they cost 2 less, but never less than 1 unit.
Under All Governance Targets, the Lab's Lingering effects setting is always on. In it, damage outlasts its cause. With it on, Store less data works at reduced strength unless Understand the system is on too. The targets can be met one way within the budget: the same five tools, with Store less data at its stronger setting, for all 14 units.
Adding Understand the system to the five instead closes every pathway for 13 units. But that level asks for the tool to be clearly helping the work, and then it helps, but not clearly.
More is not better here. Every tool at its stronger setting costs 44 units, more than three times the budget. It closes every failure pathway, but the tool then hurts the work. So it meets the targets at none of the three levels that set them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the RSF-class predictive child-safety scorer network: 5 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
The released tracking data describes a list, so the network draws that list. It holds more than 4,100 children the tool gave a 90 percent or higher chance of death or serious injury. That list was far more than any caseload could act on. Children who died in cases DCFS already knew were not on it. So one list stands for both failures: the thousands on it, and the children who died off it.
- assumed
This network follows the pattern the Illinois Rapid Safety Feedback case file describes: one tool scoring children for risk of serious harm. It does not reconstruct the actual tool.
- baseline
The pathway named One tool scores every child stands for the documented double failure. A single tool gave thousands of children extreme scores and missed children who later died.
- baseline
The documented pattern paired thousands of extreme scores with missed children at real risk. So the network assumes the score steers caseworkers' attention without reliably catching the harm it was meant to catch.
- assumed
The hotline reports are drawn as their own part of the network, because every report the tool scored brings new cases in. Mistakes can pass from the reports into the score.
- assumed
The network does not show which children and families were scored, or the harm to them. The case file records the children who died, and that scored families were typically never told they were scored.
What this example does not show
Show all 1 limitation
- This example does not show the harm to children and families, or which families bore more of it. It traces how mistakes pass between the tool, the caseworkers, and the records. The case file records the children who died. It does not show which families bore more of the harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Illinois's Rapid Safety Feedback flagged thousands of children at 90-percent-or-higher risk of serious harm — beyond any caseload's capacity to act — while children who died in known-to-system cases had not been flagged; the agency ended its use in 2017.
empirical- Investigative Chicago Tribune, Can an algorithm tell when kids are in danger? (2017) https://www.chicagotribune.com/2017/12/06/can-an-algorithm-tell-when-kids-are-in-danger/
- Investigative The Imprint, Illinois Drops Rapid Safety Feedback, A Predictive Analytics Tool (2017) https://imprintnews.org/politics/stateline-illinois-drops-rapid-safety-feedback-predictive-analytics-tool/28913
- Trade press Government Technology, Illinois Ends Child Abuse Prediction Program (2017) https://www.govtech.com/health/illinois-ends-child-abuse-prediction-program.html
Internal DCFS tracking data released under Illinois public-records law showed the Rapid Safety Feedback tool flagged more than 4,100 children at a 90-percent-or-higher probability of death or serious injury within two years, including 369 children under age 9 assigned a 100-percent probability, while children who died in cases already known to the system — among them 17-month-old Semaj Crosby, found dead after at least ten DCFS investigations — were not flagged as top-risk; the roughly $366,000 program was ended in 2017.
empirical- Investigative Chicago Tribune, Can an algorithm tell when kids are in danger? (2017) https://www.chicagotribune.com/2017/12/06/can-an-algorithm-tell-when-kids-are-in-danger/
- Trade press Government Technology, Illinois Ends Child Abuse Prediction Program (2017) https://www.govtech.com/health/illinois-ends-child-abuse-prediction-program.html
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Pause AI on alarms — Deployment circuit-breaker
- Review the riskiest first — Risk-tiered oversight
- Keep skills sharp — Deskilling-arrest mandate
- Escalate checks — State-feedback vigilance
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Require sign-off — Conformity assessment gate
- Upgrade model — Improve the model
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Store less data — Data minimization
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
Documented case histories
- Illinois Rapid Safety Feedback
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model