PAN Lab example
Allegheny Family Screening Tool
The score and the screener: a child-welfare risk tool
The Allegheny Family Screening Tool scores reports to the county's child protection hotline. Hotline staff weigh the score when deciding which families to investigate.
See more
The Allegheny Family Screening Tool is a predictive risk model run by the Allegheny County Department of Human Services in Pennsylvania since 2016. It scores each report to the child protection hotline from 1 to 20, using the county's records on the family. The score estimates how likely a child is to be placed in foster care within two years.
How it is used
Reports of possible child abuse or neglect come to the county's child protection hotline. A call screener takes each report, reads the score, and recommends whether to investigate. A supervisor makes the final decision. Sending a report for investigation is called screening it in.
County policy makes investigation the default for scores of 18 or more when a child 16 or younger is in the household. Screening such a report out needs a supervisor's explicit approval. In the county's 2019 evaluation, 61 percent of reports under the default were investigated.
Families are not allowed to know their scores, the AP reported. Erin Dalton, the county's human services director, told the AP that workers face 14,000 to 16,000 of these decisions a year, with "incredibly imperfect information."
What the score predicts
The score predicts whether a child will be placed in foster care within two years. That is a later decision by the agency itself, not the report of suspected abuse or neglect the screener must judge now. A 2026 child-welfare chapter by Zhang and Denby-Brinson treats this kind of gap as a design property of a deployment, not a flaw in the model's accuracy.
What the evaluations found
The evidence on racial fairness is contested. The county-commissioned Stanford evaluation reported that the tool and accompanying policy changes reduced racial gaps in investigation and case-opening rates. The AP reported that the developers' own unpublished analysis found no statistically significant effect of the tool on the disparity.
An independent Carnegie Mellon audit of the tool's first years, 2016 to 2018, placed the effect in the step where people overruled the score. Had the score alone decided, it would have sent about 68 percent of Black children and 50 percent of white children for investigation. Workers actually sent 51 percent and 43 percent. They disagreed with the score about a third of the time.
An ACLU and Human Rights Data Analysis Group analysis found that permanent 'ever-in' flags from public-benefits records affected 97 percent of Black households referred. The figure for non-Black households was 80 percent. The analysis cast the tool as profiling poverty and permanent records.
In 2023 the AP reported that the U.S. Department of Justice's Civil Rights Division was scrutinizing the tool. Complaints raised its use of disability, mental-health, and Supplemental Security Income data. The case file reports no public findings.
Two loops to watch
The first loop is memory. Today's screening decisions are recorded in the county's records, and tomorrow's scores are computed from those records.
The second loop is deference. Screeners and supervisors can come to lean on the score for the reports where they have a choice. That reliance adds to the default the policy already sets for the highest scores.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. Here a mistake is, for example, a score that overstates a family's risk, or a wrong entry in a family's records. A pathway counts as closed once it passes on only a few mistakes. It need not stop them all.
This case has a budget of 9 units. Explore (No Targets) sets no targets. There, and under Service Targets Only, one tool costing 2 units is enough for the network to correct its mistakes rather than build on them. Escalate checks does it, and so does Mark AI-written records. Under Service Targets Only, either one also meets the service target. That target asks that the score keep clearly helping the screening work.
Under Service and Safety Targets and All Governance Targets, the targets can be met. Both levels also ask you to close every failure pathway. Four are open at the start: the score shown to screeners, screening decisions recorded, past records used in each score, and records read at screening.
Five sets of tools meet the targets at each of these two levels. Every one includes Escalate checks, Mark AI-written records, and Store less data, which together cost 7 units. Escalate checks closes the score shown to screeners. Mark AI-written records closes the two pathways out of the records. Store less data closes screening decisions recorded. The other four sets add one more tool: Keep skills sharp, Keep prompts neutral, Review on schedule, or Assign a challenger. In the set of three, any one tool may be at its stronger setting instead, within the budget.
In this network the score writes nothing into the records directly. Here Mark AI-written records marks which past decisions followed the score. This tool takes no prompts, so here Keep prompts neutral acts on which reports get scored. Lingering effects is a Lab setting in which effects last after their cause is gone. With lingering effects on, as under All Governance Targets, Store less data and Review the riskiest first work at reduced strength unless Understand the system is also on. Store less data still closes its pathway.
More is not better here. Every tool at its strongest setting at once closes every failure pathway, but costs 32 units, far over the budget. It also leaves the score no longer clearly helping the screening work. So it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the AFST-class human-in-the-loop risk tool network: 4 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
This network follows the pattern of a risk score that people review, as documented in the Allegheny Family Screening Tool case file. It does not rebuild the actual tool.
- assumed
The network assumes screening staff affect one another in two ways. Screeners compare scores and calls informally, which can spread reliance on the score. Supervisors' second reads of screeners' recommendations, which the county documents, can catch mistakes.
- baseline
The network assumes, from the start, that screening decisions feed into later scores. The case file says the score is computed from administrative records. Records of earlier reports and decisions become inputs to later scores.
- assumed
The screening supervisors have their own pathway, based on the supervisory review the county documents. So their part counts in how mistakes pass. Whether that review reduced the documented racial disparity is recorded in the case file, not computed here.
- assumed
The network assumes screeners use real judgment, but the score pulls on it. The score shifts their judgment rather than replacing it.
What this example does not show
Show all 1 limitation
- Bias moves here the way mistakes do: through how reports are framed, what is recorded, and what is read back. The Lab models no demographics. It estimates no unequal harm to the families served. That harm is documented in the case files and measured outside any network like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.
empirical- Government evaluation Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf
- Academic Centre for Social Data Analytics (AUT), AFST evaluation summary https://csda.aut.ac.nz/news-and-events/2019/allegheny-family-screening-tool-evaluation-improved-decision-accuracy,-reduced-disparities
- Academic Rittenhouse, Algorithms, Humans and Racial Disparities in Child Protective Services https://krittenh.github.io/katherine-rittenhouse.com/Rittenhouse_Algorithms.pdf
In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.
empirical- Academic Rittenhouse, Algorithms, Humans and Racial Disparities in Child Protective Services https://krittenh.github.io/katherine-rittenhouse.com/Rittenhouse_Algorithms.pdf
- Government evaluation Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf
- Academic Centre for Social Data Analytics (AUT), AFST evaluation summary https://csda.aut.ac.nz/news-and-events/2019/allegheny-family-screening-tool-evaluation-improved-decision-accuracy,-reduced-disparities
- Peer-reviewed Stapleton, L., Lee, M. H., Qing, D., Wright, M., Chouldechova, A., Holstein, K., Wu, Z. S., & Zhu, H. (2022). Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders. 2022 ACM Conference on Fairness Accountability and Transparency, 1162–1177. https://doi.org/10.1145/3531146.3533177
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Keep prompts neutral — Framing and mirroring reduction
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Mark AI-written records — Provenance labeling
- Upgrade model — Improve the model
- Store less data — Data minimization
- Assign a challenger — Structured dissent
Documented case histories
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model