Skip to content

PAN Lab example

Illinois Rapid Safety Feedback

The alarm that cried wolf: a child-safety scorer

Illinois's child-welfare agency scored children in reports of suspected abuse or neglect. Its tool flagged thousands as extreme risks and missed children who died.

See more

Rapid Safety Feedback was a predictive tool from the nonprofit Eckerd Connects and its for-profit partner, MindShare Technology. Illinois's Department of Children and Family Services, DCFS, used it until December 2017. Each night it scored every child named in a hotline report of suspected abuse or neglect for the chance of death or serious injury within two years.

How it was meant to work

Each score ran from 1 to 100. A high score was designed to prompt a supervisor's review, a "second set of eyes", and coaching for the caseworker.

Eckerd says front-line caseworkers should never get the raw scores, let alone decide on them. It says DCFS supervisors, trained and coached by Eckerd, should review the scores and decide which cases need immediate attention.

What went wrong

DCFS's internal tracking data, released under Illinois public-records law, showed more than 4,100 children scored at 90 percent or higher. That included 369 children under age 9 given a score of 100 percent. It was far more than any caseload could act on. Reporting found caseworkers alarmed and overwhelmed by the alerts.

Meanwhile, children who died in cases DCFS already knew were not flagged as top risk. Semaj Crosby, 17 months old, was found dead after at least ten DCFS investigations. Itachi Boyle, 22 months old, also died without a high score.

The records behind the score

The scores were recomputed each night from DCFS's records, and those records were flawed. They had entry errors. They often failed to link a child's history to siblings or other adults in the home. State law required DCFS to erase investigations closed as "unfounded".

So the tool worked from a patchy, incomplete picture. It scored thousands of children as extreme risks while missing children with long histories with DCFS.

Why the alarm failed both ways

A score is calibrated when it means what it says: a 30 percent score should match about 30 cases in 100. The case file says this failure ran deeper than calibration, because the records behind the score were patchy.

A flag that fires for thousands cannot change where scarce attention goes. It wears people down with alarms. It also makes an unflagged case look cleared. Trusting such an alarm brings false panic and false calm together.

How it came in and how it ended

George Sheldon became DCFS director in 2015, after a run of child deaths. The roughly $366,000 program was central to the reforms he promised. DCFS brought it in through a no-bid arrangement that the state classified as a grant.

In July 2017, a joint report by the Illinois Executive Inspector General and the DCFS Inspector General found that classification was mismanagement. It had sidestepped the state's rules for transparent bidding.

In December 2017 the new director, Beverly "B.J." Walker, ended the predictive scoring. She said it "wasn't predicting any of the bad cases." DCFS kept a smaller Eckerd case-review training program. Walker said 15 staff and three supervisors were using it.

Who caught it

The decisive corrections came from outside DCFS. They were investigative reporting, a public-records release that put the extreme scores in plain view, and the inspectors general's finding on how the program was bought.

Beyond Illinois

Reporting and a later review by the American Civil Liberties Union placed Illinois among several jurisdictions that took up the tool and then dropped it. Alaska, Louisiana, Ohio, and Oklahoma were among them. The scored families were typically never told they were scored.

What this network is drawn from

This network follows the pattern the case file describes. It is not a reconstruction of the actual tool. It shows the scorer, the hotline reports, the list of highest scores, the DCFS caseworkers and supervisors, and DCFS's case records.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.

This case has a budget of 14 units. Before any tool is used, mistakes build on one another across the network, and eight failure pathways are open.

Explore (No Targets) sets no targets. There, and under Service Targets Only, one tool is enough to stop mistakes building on one another: Mark AI-written records, at 2 units. Under Service Targets Only it also meets the service target, which asks that the tool be helping the work. Many other combinations do both.

Service and Safety Targets also asks you to close every failure pathway, among other targets. The targets can be met within the budget. The cheapest way costs 12 units. It uses five tools: Escalate checks, Mark AI-written records, Keep prompts neutral, Store less data, and Check with a second model.

Every way to meet these targets includes all five. Counting each tool's stronger setting as a separate choice, there are 25 ways within the budget. They use 7 different sets of tools.

Understand the system costs 4 units at the two upper levels and 3 units at the two lower ones. Its stronger setting costs 6. While it is on, Mark AI-written records, Store less data, Check with a second model, and Require sign-off each cost 1 unit less. At its stronger setting they cost 2 less, but never less than 1 unit.

Under All Governance Targets, the Lab's Lingering effects setting is always on. In it, damage outlasts its cause. With it on, Store less data works at reduced strength unless Understand the system is on too. The targets can be met one way within the budget: the same five tools, with Store less data at its stronger setting, for all 14 units.

Adding Understand the system to the five instead closes every pathway for 13 units. But that level asks for the tool to be clearly helping the work, and then it helps, but not clearly.

More is not better here. Every tool at its stronger setting costs 44 units, more than three times the budget. It closes every failure pathway, but the tool then hurts the work. So it meets the targets at none of the three levels that set them.

Stylized model of a documented deploymentChild welfare & family services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the RSF-class predictive child-safety scorer network: 5 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 6 assumptions
  • assumed

    The released tracking data describes a list, so the network draws that list. It holds more than 4,100 children the tool gave a 90 percent or higher chance of death or serious injury. That list was far more than any caseload could act on. Children who died in cases DCFS already knew were not on it. So one list stands for both failures: the thousands on it, and the children who died off it.

  • assumed

    This network follows the pattern the Illinois Rapid Safety Feedback case file describes: one tool scoring children for risk of serious harm. It does not reconstruct the actual tool.

  • baseline

    The pathway named One tool scores every child stands for the documented double failure. A single tool gave thousands of children extreme scores and missed children who later died.

  • baseline

    The documented pattern paired thousands of extreme scores with missed children at real risk. So the network assumes the score steers caseworkers' attention without reliably catching the harm it was meant to catch.

  • assumed

    The hotline reports are drawn as their own part of the network, because every report the tool scored brings new cases in. Mistakes can pass from the reports into the score.

  • assumed

    The network does not show which children and families were scored, or the harm to them. The case file records the children who died, and that scored families were typically never told they were scored.

What this example does not show

Show all 1 limitation
  • This example does not show the harm to children and families, or which families bore more of it. It traces how mistakes pass between the tool, the caseworkers, and the records. The case file records the children who died. It does not show which families bore more of the harm.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Illinois's Rapid Safety Feedback flagged thousands of children at 90-percent-or-higher risk of serious harm — beyond any caseload's capacity to act — while children who died in known-to-system cases had not been flagged; the agency ended its use in 2017.

    empirical
    • Investigative Chicago Tribune, Can an algorithm tell when kids are in danger? (2017) https://www.chicagotribune.com/2017/12/06/can-an-algorithm-tell-when-kids-are-in-danger/
    • Investigative The Imprint, Illinois Drops Rapid Safety Feedback, A Predictive Analytics Tool (2017) https://imprintnews.org/politics/stateline-illinois-drops-rapid-safety-feedback-predictive-analytics-tool/28913
    • Trade press Government Technology, Illinois Ends Child Abuse Prediction Program (2017) https://www.govtech.com/health/illinois-ends-child-abuse-prediction-program.html
  • Internal DCFS tracking data released under Illinois public-records law showed the Rapid Safety Feedback tool flagged more than 4,100 children at a 90-percent-or-higher probability of death or serious injury within two years, including 369 children under age 9 assigned a 100-percent probability, while children who died in cases already known to the system — among them 17-month-old Semaj Crosby, found dead after at least ten DCFS investigations — were not flagged as top-risk; the roughly $366,000 program was ended in 2017.

    empirical
    • Investigative Chicago Tribune, Can an algorithm tell when kids are in danger? (2017) https://www.chicagotribune.com/2017/12/06/can-an-algorithm-tell-when-kids-are-in-danger/
    • Trade press Government Technology, Illinois Ends Child Abuse Prediction Program (2017) https://www.govtech.com/health/illinois-ends-child-abuse-prediction-program.html

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Child welfare & family services domain page.

Levers available here and the patterns behind them

Documented case histories