PAN Lab example
Douglas County Decision Aide
The score read only at the edges: a child-welfare screening aide
Douglas County, Colorado's Decision Aide scores reports of possible child maltreatment (referrals). An independent trial found it sped decisions without significantly changing children's outcomes.
See more
The Douglas County Decision Aide is a risk score used by the Department of Human Services in Douglas County, Colorado. It rates each report of possible child maltreatment from 1 to 20 for how likely the child is to be removed from home within two years. It is computed from state child-welfare, public-benefit, and court records.
How it is used
Screening is the RED Team's decision on each referral: whether to screen it in or screen it out. The sources do not say what follows each choice.
The score is built into the county's RED Team screening. RED stands for Read, Evaluate, Direct. A supervisor and at least two caseworkers decide together how to screen each referral. The county uses these teams for roughly 85 percent of referrals.
The team may consult the score, but consulting it is voluntary, and the team keeps the decision. County officials said they would share a family's score on request.
How it was built
Douglas County's Department of Human Services commissioned the Centre for Social Data Analytics at Auckland University of Technology in 2017. A feasibility stage was completed that August.
The Decision Aide is a LASSO logistic regression, a statistical model that drops the predictors that add little to its estimate. Race-related predictors were tested and deliberately left out of the model in use. The developers report that it is well calibrated across race, age, and gender groups. Each quarter, the developers also checked whether its accuracy was shifting over time. These are the drift checks.
What the trial found
The Decision Aide launched in February 2019 as a year-long randomized controlled trial. Whole teams were randomly assigned either to use it or to screen as they always had. Cornell researchers Maria Fitzpatrick and Christopher Wildeman, with Katharine Sadowski, evaluated it independently.
They used an ethics framework carried over from an evaluation in Allegheny County, Pennsylvania, which runs its own child-welfare screening tool.
The peer-reviewed trial found the Decision Aide sped up screening decisions without significantly changing outcomes for children. COVID limited the outcome analysis. A companion study found workers attended mainly to extreme high or low scores. They largely disregarded mid-range ones.
What the case file draws from it
On paper, this is the most carefully governed deployment the case file covers. There was an independent randomized trial, an ethics review, and ongoing drift monitoring. Race predictors were left out, and a consensus of several people was required. County officials said they would share scores with families on request.
Still, the trial found only a modest effect. The case file's reading is that the limit was not the model's accuracy. It was how people used the score. Workers reacted to the extremes and rarely moved mid-range decisions.
The step from a score to a next action was the one thing left ungoverned, and it decided the impact. A safeguard as loose as "the team may consider the score" can leave a score with almost no effect, in either direction.
Which evidence belongs to this tool
The press often credits this tool with strong positive results: a large drop in child-injury hospitalizations and reduced surveillance of low-risk Black children. Those results come from a separate deployment in another Colorado county, not from Douglas County. The case file calls keeping straight which evidence belongs to which system a governance discipline in itself.
National reporting later grouped Douglas County's tool with others inspired by Allegheny County's. Colorado headlines tied it to a U.S. Justice Department inquiry. That inquiry targeted Allegheny County's tool, not Douglas County's.
No comparison of Douglas outcomes by race was measured. A public-records request filed through MuckRock found the county held no such data.
What this network is drawn from
This network follows the pattern the case file documents. It does not reconstruct the actual tool. It shows the Decision Aide, the RED Team, the state records, and the independent trial and drift checks.
Watch two things. First, a loop: today's screening decisions become part of the records tomorrow's scores are computed from. Second, the team's discretion does more of the work than the score does.
Where the facts come from
The sources include the developers' 2019 methodology report and their project page. They include the peer-reviewed trial in the Journal of Human Resources (2025) and the companion study in Sociological Science (2026). They also include the trial's public registration, Associated Press reporting from 2022 and 2023, and the MuckRock public-records request.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 9 units. Understand the system is a tool that funds ongoing study of what the deployment is really doing. It costs 3 units under Explore (No Targets) and Service Targets Only, and 4 under the two higher levels. Its stronger setting costs 6. While it is on, four tools cost 1 unit less: Keep skills sharp, Review the riskiest first, Mark AI-written records, and Review on schedule. At its stronger setting they cost 2 less, never below 1.
Lingering effects is a Lab setting in which mistakes and reliance on the system stay after their cause is gone. It starts on under Explore (No Targets) and Service Targets Only, where you can switch it off. Service and Safety Targets keeps it off, and All Governance Targets keeps it on. While it is on, three tools work at reduced strength unless Understand the system is also on. They are Review the riskiest first, Store less data, and Vet connections. Vet connections then closes no pathway on this network.
Explore (No Targets) sets no targets. Service Targets Only asks for two things. The network must be self-correcting, meaning its mistakes are corrected rather than building on each other. The Decision Aide must also be helping the screening work. Before any tool is used, it is helping, but the network is at a tipping point, not self-correcting. At a tipping point, the network is neither clearly correcting its mistakes nor letting them cascade.
Seven tools meet Service Targets Only on their own at their standard setting. Keep skills sharp, Mark AI-written records, Review on schedule, and Escalate checks cost 2 units each. Review the riskiest first, Understand the system, and Upgrade model cost 3. Store less data does it at its stronger setting, for 5 units.
With lingering effects off, Store less data and Vet connections also do it alone, for 3 units each. Upgrade model then needs its stronger setting, for 5. Keep prompts neutral never does it alone. With lingering effects on, 285 different sets of tools within the budget meet this level.
Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, among other targets. Four are open before any tool is used. They are Score given to the RED Team, Screening decision written to the records, Records used to compute each score, and Team reads prior history.
The cheapest way costs 7 units, and each of its three tools closes its own part. Escalate checks closes Score given to the RED Team. Mark AI-written records marks entries in the state records that a machine produced, such as logged scores. The RED Team and the Decision Aide can then weigh them accordingly. It closes the two pathways that read from the records.
One of them, Records used to compute each score, is also the pathway that lowers the Privacy gauge. The Privacy gauge shows how well personal information in the network is protected. Store less data closes Screening decision written to the records.
Four different sets of tools meet each of these levels. One is those three tools. Each of the other three adds Review on schedule, Keep prompts neutral, or Keep skills sharp, for 9 units. Three more ways use one of the three at its stronger setting: Mark AI-written records for 8 units, or Escalate checks or Store less data for 9. Review the riskiest first, Upgrade model, Vet connections, and Understand the system are in no way that meets either level.
More is not better here. Every tool at its strongest setting costs 32 units, more than three times the budget. It makes the network self-correcting and closes every failure pathway. But the Decision Aide then adds too little to the screening work, so it meets no level's targets.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Douglas-class independently-evaluated screening aide network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
This network follows the pattern the Douglas County Decision Aide case file documents. It is a screening score that was independently evaluated, used by a team that keeps real discretion. It does not reconstruct the actual tool.
- assumed
The network draws four links that start and end at the same part. The RED Team's consensus is a documented, built-in check: a supervisor and at least two caseworkers must agree, so each member's reading is checked against the others'. The network also assumes team members pull one another toward a shared view while reaching consensus. One model scores every referral, so its blind spots are shared across all of them. The fourth link is a check of each score by a second, independent model, which the sources do not describe.
- baseline
The network assumes the team keeps real discretion. The independent trial found the score sped decisions without significantly changing outcomes for children. A companion study found workers attended mainly to extreme scores. So the score anchors the decision rather than replacing it.
- assumed
The network assumes screening decisions shape future scores. Scores are computed from accumulated records from several agencies. The outcome they predict, removal from home, is itself the result of earlier human decisions. So today's screening shapes tomorrow's inputs.
- assumed
Race-related predictors were deliberately left out of the model. No comparison of Douglas outcomes by race was measured, because the county held no such data. The strong positive findings often credited to this tool come from a separate deployment in another Colorado county. This network shows how mistakes pass between the tool, the team, and the records. It does not model demographics, and it estimates no unequal harm to the families served.
What this example does not show
Show all 2 limitations
- This example does not show demographics or unequal harm to the families served. The Decision Aide left out race-related predictors, and no comparison of its outcomes by race was measured. The example shows how mistakes pass between the tool, the team, and the records.
- This example does not show the strong positive results often credited to this tool in the press. Those are a large drop in child-injury hospitalizations and reduced surveillance of low-risk Black children. They come from a separate deployment in another Colorado county, documented in the case file, not from Douglas County.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019, scores each referral from 1 to 20 for a child's likelihood of out-of-home removal within two years; an independent Cornell-led randomized controlled trial found it sped up screening decisions without significantly changing child outcomes, and a companion study found workers attended mainly to extreme scores while largely disregarding mid-range ones.
empirical- Academic Vaithianathan et al. (Centre for Social Data Analytics, AUT), Implementing a Child Welfare Decision Aide in Douglas County: Methodology Report (2019) https://csda.aut.ac.nz/__data/assets/pdf_file/0009/347715/Douglas-County-Methodology_Final_3_02_2020.pdf
- Academic Fitzpatrick, Sadowski and Wildeman, Algorithms and Decision-making: Evidence from Child Maltreatment Reports (Journal of Human Resources, 2025) https://jhr.uwpress.org/content/early/2025/08/01/jhr.0224-13437R2
- Academic Eiermann, Fitzpatrick, Sadowski and Wildeman, How Do (Human) Child Welfare Workers Respond to Machine-Generated Risk Scores? (Sociological Science, 2026) https://sociologicalscience.com/articles-v13-1-1/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Keep skills sharp — Deskilling-arrest mandate
- Keep prompts neutral — Framing and mirroring reduction
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Store less data — Data minimization
- Upgrade model — Improve the model
- Escalate checks — State-feedback vigilance
- Vet connections — Connection authorization
Documented case histories
- Douglas County Decision Aide
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model