PAN Lab example
Oregon Safety at Screening
The fix then the off switch: a fairness-corrected screening tool
Oregon's child-welfare hotline screeners saw a fairness-corrected risk score before deciding whether to investigate a report. The agency stopped using the tool in 2022.
See more
The Safety at Screening Tool was software that Oregon's Department of Human Services built in house through its analytics office. For each child named in a report of abuse or neglect, it estimated two risks from the state's own child-welfare records. Hotline screeners saw the scores after entering a report and before deciding whether to investigate.
What the scores estimate
One score estimated the chance a child would be removed from home within two years if the report was assigned for investigation. The other applied if the report was screened out, meaning closed without an investigation. It estimated the chance the child would later be named in a new report that was assigned for investigation.
Two separate models produced them. One was trained on past reports that were investigated, and one on past reports that were not. The agency split them because the two groups of children can go on to very different experiences.
Each probability was shown as a tier from 1 to 4. Tier 1 covered the lowest tenth of probabilities, and tier 4 the top 5 percent. The screen also showed a report score, the highest score among the children on the report. Children already in substitute care, such as foster care, were not scored. Reports about these children made up 4 percent of incoming reports.
What the scores are built from
Both models used only the agency's own records in OR-KIDS, Oregon's child-welfare case system, online since August 2011. No call text or voice was used. The models learned from hundreds of thousands of past reports.
The agency's report lists more than 180 variables. They cover past reports, investigations, and foster care, the current report's details, the alleged perpetrator, and who made the report. The list includes the child's race or ethnicity category.
The agency's view was that other data can stand in for race, so hiding it would not remove bias. It chose to address bias directly instead.
The fairness correction
After the models ran, Oregon applied a fairness correction. It set different tier cut-offs for different race and ethnicity groups. The goal was error rate balance: similar rates of two errors across groups.
One error rates a child high risk when the predicted event, a removal or a new investigated report, does not happen. The other rates a child low risk when it does happen. The agency's 2019 report says the correction raised its error rate balance measure from 0.43 to 0.76 on past data. A measure of 1 would mean identical error rates across groups. Accuracy of the removal model fell from 81 to 79 percent.
The sources read do not report the correction's effect once the tool was in use.
Who decides
Hotline screeners are trained state employees who take reports of abuse or neglect, usually by phone. They decide whether to screen a report out or assign it for investigation by Child Protective Services.
They saw the scores only after entering the report, so the scores would not shape what they gathered. They affirmed they had reviewed the scores, then decided. The decision stayed theirs.
The agency framed the scores as historical indicators, not answers. It used four tiers so screeners would not set their own cut-offs. It also planned to track data entry, to check that it stayed consistent with past patterns. One concern it named was screeners changing entries to get a desired score.
Why the agency built it
In 2018, screeners assigned 51 percent of 85,974 reports for investigation, up from 43 percent of 67,466 reports in 2012. The agency's 2019 report says that volume strained Child Protective Services.
It also found that, over two years, 13 percent of assigned reports led to a child's removal from home. So did 8 percent of reports that were screened out.
The tool was derived from the Allegheny Family Screening Tool, a screening tool used in Allegheny County, Pennsylvania. Oregon narrowed its version to the agency's own child-welfare records.
How it ended
The tool was first used in 2018. On 19 May 2022, deputy director Lacey Andresen emailed staff. After what the agency called “extensive analysis”, hotline workers would stop using the tool at the end of June. “We are committed to continuous quality improvement and equity,” she wrote. The Associated Press reported the aim was to reduce disparities in which families are investigated.
The agency replaced it with Structured Decision Making, a screening process with no algorithm. Spokesman Jake Sunderland said the tool's score, built from aggregate data, could not be used in that family-specific process.
What each side says
The decision came weeks after Associated Press reporting on the Allegheny tool. That reporting cited a Carnegie Mellon study. It found that if the Allegheny tool had decided on its own, with call volumes held comparable, it would have recommended investigating about two-thirds of Black children. For all other children, it was about half.
U.S. Senator Ron Wyden of Oregon then asked the agency again about racial bias. He said such decisions are “far too important a task to give untested algorithms.”
Sunderland told Willamette Week the decision was unrelated to the coverage. He said the agency chose in 2021 to replace its screening process. Oregon's own tool was never publicly shown to reproduce the Allegheny disparities.
What this case shows
The score is computed from records that earlier screening decisions helped write. The fairness correction was built to offset the bias in those records, not to cut that loop.
The agency also held, and used, the authority to stop the tool. The case file calls this a rare case: the agency stopped its own tool before a documented crisis forced it. The most heavily documented cases, it notes, ended in courts or commissions. It also notes that the two stated reasons, equity and a poor fit with the new process, sit in tension.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it goes on.
This case has a budget of 9 units. Explore (No Targets) sets no targets. There, the cheapest single tools that keep mistakes from building on one another across the network are Escalate checks and Mark AI-written records, at 2 units each. Several others do it alone at a higher cost.
Under Service Targets Only, each of them also meets the service target. That target asks that the score stay helpful to the screeners' work. No combination that includes Pause AI on alarms meets it, because halting the score leaves it no longer helping.
Under Service and Safety Targets and All Governance Targets, the targets can be met. Both levels also ask you to close every failure pathway, among other targets. The cheapest way costs 7 of the 9 units: Escalate checks, Mark AI-written records, and Store less data.
Every combination that meets these targets includes those three. Counting sets that need a stronger setting, five sets of tools within the budget meet them at each of the two levels.
Escalate checks closes Scores shown to screeners. Mark AI-written records closes Past records used in scoring and Case history read by screeners. Store less data closes Screening decisions entered in OR-KIDS.
Understand the system costs 3 units, or 4 under the two higher levels, and 6 at its stronger setting. While it is on, four tools cost less, never below 1 unit: Review the riskiest first, Keep skills sharp, Pause AI on alarms, and Require sign-off. It is in no combination that meets the two higher levels.
Lingering effects is a Dynamics setting in which damage outlasts its cause. Under All Governance Targets, lingering effects are always on. Then Store less data, Vet connections, and Review the riskiest first work at reduced strength unless Understand the system is also on. Store less data still closes its pathway there.
More is not better here. Every tool at its strongest setting at once closes every failure pathway, but costs 39 units, far over the budget. It also leaves the score no longer helping, so it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Oregon-class fairness-corrected screening tool network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 6 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
This example follows the fairness-corrected screening tool that Oregon later stopped, as documented in the case file on Oregon's Safety at Screening Tool. It does not rebuild the real tool.
- assumed
This example places the fairness correction between the score and the screeners, where the agency's report puts it: after the models run, before screeners see a tier. It has no pathways of its own, so it changes nothing in how mistakes move. The agency's 2019 report measured its effect on past data. The sources read do not report its effect once in use.
- baseline
This example includes Oregon DHS leadership as a reviewer that receives the score's results. It stands for the agency's authority to stop the tool, which it used in 2022. That was a single decision, not a standing review. The sources do not describe what the agency's analysis examined.
- assumed
This example assumes earlier screening decisions count in later scores from the start. The agency's 2019 report says its administrative data holds years of human data entry and decisions, so it likely contains inaccuracies and bias. The fairness correction was built to offset that bias, not to cut the loop.
- assumed
Oregon limited the models to its own child-welfare records, with no call text or voice. It also built in guards against over-reliance: four coarse tiers, scores shown only after data entry, and scores framed as historical indicators. So this example treats the score as one input to the screener's decision, not the deciding one. It does not count the agency's reuse of its own records as a privacy exposure.
- assumed
Screeners kept real discretion. The score was advisory, and screeners had only to affirm they had reviewed it. The decision to screen a report out or investigate it stayed theirs.
- assumed
The case file records concerns about racial disparity in child-welfare screening, and the agency's report measured the fairness correction's effect on error rates. Both concern children and families. This example shows how errors move among the score, the screeners, and the records. It does not model groups of people or estimate unequal harm to them.
What this example does not show
Show all 1 limitation
- This example shows how mistakes move through the records, the screeners who read them, and the score that counts earlier decisions. It does not model groups of people or estimate unequal harm to children and families. The case file records racial-disparity concerns about the Allegheny tool Oregon's was derived from. Oregon's own tool was never publicly shown to reproduce them. The agency's report measured the fairness correction's effect on past data, outside any network like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Oregon's child-welfare agency dropped its AFST-derived Safety at Screening tool in 2022, citing equity concerns amid national scrutiny of racial disparity in child-welfare algorithms.
empirical- Investigative NPR/AP, Oregon is dropping an AI tool used in child welfare system (2022) https://www.npr.org/2022/06/02/1102661376/oregon-drops-artificial-intelligence-child-abuse-cases
- Investigative Willamette Week, Oregon DHS to End Its Use of Child Abuse Risk Algorithm (2022) https://www.wweek.com/news/state/2022/06/04/oregon-department-of-human-services-ends-its-use-of-child-abuse-risk-algorithm/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Escalate checks — State-feedback vigilance
- Pause AI on alarms — Deployment circuit-breaker
- Require sign-off — Conformity assessment gate
- Review on schedule — Oversight cadence & retrospectives
- Keep skills sharp — Deskilling-arrest mandate
- Mark AI-written records — Provenance labeling
- Store less data — Data minimization
- Vet connections — Connection authorization
- Review the riskiest first — Risk-tiered oversight
- Understand the system — Understand the system
- Upgrade model — Improve the model
- Assign a challenger — Structured dissent
Documented case histories
- Oregon Safety at Screening
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model