PAN Lab example
Los Angeles County Project AURA
Caught at the gate: a child-abuse risk model that never shipped
Los Angeles County tested Project AURA, a child-fatality risk score, on past cases. It falsely flagged 3,829 children, correctly flagged 171, and was never used.
See more
Project AURA was a predictive risk model that the analytics firm SAS built for Los Angeles County's Department of Children and Family Services (DCFS). It scored each child in a referral, a report of suspected abuse or neglect, from 1 to 1,000, using records from several agencies. How SAS's model produced each score was proprietary.
What the test found
AURA was trained and tested on past cases, never on a live one. The test compared its flags with child deaths, near-fatalities, and critical incidents from 2011 and 2012. The sources do not define a critical incident.
The sources report results at one high-risk cut: the score above which AURA flagged a child. At that cut, AURA correctly flagged 171 children who had one of those worst outcomes. It also flagged 3,829 children who did not. Those are false flags, or false positives. That is about 22 false flags for every correct one.
DCFS's own public-affairs director confirmed on the record a false-positive rate of no less than 95.6 percent. That is the share of flagged children who did not have one of those outcomes.
What else the sources say about the model
The algorithm was proprietary to SAS, and no one outside SAS knows exactly how it works. WitnessLA says SAS ran the check of what actually happened to the families.
How it ended
AURA was never used on a single live case. In May 2017 the county's Office of Child Protection reported to the Board of Supervisors that DCFS was no longer pursuing it. By then the false-positive results had been known for nearly two years.
A 2017 commentary by Richard Wexler in WitnessLA says that news was buried on page 10 of the report. It says it is not clear what finally prompted DCFS to drop AURA. Perhaps, it says, it was the report's point that all those false positives would further overload the system.
More likely, the commentary says, it was a State of California initiative to come up with a "better" predictive analytics model.
The same report called for strict standards before considering the use of predictive-analytics models. They include understanding how racism and other biases may be embedded in systemic data.
The background
DCFS contracted SAS around 2013. AURA's development unfolded amid intense scrutiny of DCFS after the May 2013 death of eight-year-old Gabriel Fernandez in Palmdale.
What came after
In 2021 the county launched a separate, in-house tool, the Risk Stratification Pilot. Unlike AURA, it is not proprietary, and it draws only on the county's own records. It deliberately excludes race, ethnicity, and geography.
Why this case matters
The case file reads AURA's test as exposing a base-rate problem that no accuracy figure could hide. A base rate is how common an outcome is. When the outcome is rare, even a fairly accurate screener flags many children who will not have it.
The county was evaluating a model it could not open. What it had was an outcome count, not a look at how the model worked. The decision also came slowly.
The case file reads refusing deployment as a governance decision in its own right, not a failure to ship. The question this case asks is not how to run the tool safely. It is which controls decide, before go-live, whether it runs at all.
Where the facts come from
The sources include 2015 reporting in The Imprint, 2015 reporting by KPCC/LAist, and a 2015 Child Protective Services Defense article. They also include Richard Wexler's 2017 commentary on the NCCPR blog and in WitnessLA. The Wexler pieces and the Child Protective Services Defense article are advocacy.
Later sources cover the Risk Stratification Pilot, from the county, DCFS, and the Children's Data Network at the University of Southern California. An NBC Los Angeles timeline covers the Gabriel Fernandez case.
What this network is drawn from
This network follows the pattern the sources describe. It is not a reconstruction of the actual tool. AURA never ran on a live case, so the network shows the shape it would have had in use.
The retrospective test used de-identified cross-agency records, with the details that identify a person removed.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.
This case has a budget of 10 units. Before any tool is used, six failure pathways are open. The network is at a tipping point, where mistakes are on the edge of building on one another.
Explore (No Targets) sets no targets. There, and under Service Targets Only, a single tool is enough to keep mistakes from building on one another. Three do it for 2 units each: Assign a challenger, Escalate checks, and Mark AI-written records. Understand the system also does it alone, for 3 units. So does Upgrade model at its stronger setting, for 5 units, with lingering effects on.
Under Service Targets Only, each of those also meets the service target. That target asks for the risk score to be helping the investigators' work. Hundreds of other combinations within the budget do too.
Under Service and Safety Targets and All Governance Targets, the targets can be met in exactly one way within the budget. Both levels also ask you to close every failure pathway. The one way uses four tools at their standard settings, for all 10 units. They are Mark AI-written records, Escalate checks, Check with a second model, and Store less data.
Each of the four closes pathways the others leave open. Mark AI-written records labels the scores logged in the records. Mistakes then stop passing from the records and the cross-agency data into the score, and from the records to investigators.
Escalate checks has investigators check the score more closely when monitoring flags trouble. Mistakes then stop passing along Score to investigators. Store less data stops them passing along Investigation findings to the records. Check with a second model stops them passing along One model for every referral.
With the budget set aside, every combination that meets these two levels still includes all four. At All Governance Targets, Store less data works at reduced strength unless Understand the system is also on. It still does enough here.
The tools closest to the county's decision not to deploy AURA close no pathway here. Require sign-off, Gate vendor updates, and Review on schedule change none. Assign a challenger keeps mistakes from building but closes none. Pause AI on alarms stops investigators getting the score, but the work then falls behind.
This network shows the score already in use. So no tool here takes it out of the investigations the way the county's decision did.
More is not better here. Every tool at its strongest setting at once costs 41 units, about four times the budget. It closes every failure pathway and keeps mistakes from building. But the risk score then no longer helps the investigators' work enough, so this misses the targets at all three levels that set them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the AURA-class pre-deployment risk scorer network: 5 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 6 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
The retrospective test is this case's main event, so the network includes it as a part of its own. It compared AURA's flags with past outcomes and found 171 correct flags and 3,829 false ones. The network includes both of its pathways, because both happened: the test read the records, and the county never deployed AURA. A test before any live use is what sets this case apart from deployments that found their problems in use.
- assumed
This network follows the pattern the Los Angeles County Project AURA case file describes: a risk model tested before any use. It is not a reconstruction of the actual tool.
- assumed
AURA was tested only against past cases and never used on a single live referral. This network shows the shape it would have had in use, with the score weighing heavily in investigation decisions. That is a choice made in building the network, not a record of live use.
- baseline
The retrospective test found 171 correct flags and 3,829 false ones. The county's Department of Children and Family Services (DCFS) put the false-positive rate, the share of flags that were false, at no less than 95.6 percent. The network assumes the score would steer investigators' attention while producing more false flags than they could work through.
- assumed
The network includes pathways where one investigator's habits pass to another and where one model's errors repeat. It also includes two checks the sources never describe happening: a second model checking AURA's scores, and a colleague challenging its rankings. Tools can add them.
- assumed
The network includes the pathway from the accumulated cross-agency records into the score. The case documents describe a model computed from administrative history gathered for other purposes.
- assumed
The harm the sources document is the number of false flags, not a measured difference between groups of families. The model's inputs, such as prior referrals, say nothing by themselves about whether it treated families fairly. This Lab shows how mistakes pass between the parts of a deployment. It does not estimate harm to particular groups of people.
What this example does not show
Show all 2 limitations
- AURA was never used on a live case. This example shows the shape it would have had in use. So the harm it shows is harm the tool would have caused, not harm that happened. It also shows the test that came before any live use.
- The harm the sources document is the number of false flags, not a measured difference between groups of families. The model's inputs, such as prior referrals, say nothing by themselves about whether it treated families fairly. This Lab shows only how mistakes pass between the parts of a deployment. Any harm to children and families is documented in the case file and measured outside any network like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built by SAS — correctly flagged 171 of the highest-risk children but produced 3,829 false positives, a false-positive rate of no less than 95.6% that DCFS's own public-affairs director confirmed on the record, and the county shelved the tool in 2017 without ever using it on a live case.
empirical- Investigative The Imprint (Daniel Heimpel), Uncharted Waters: Data Analytics and Child Protection in Los Angeles (2015) https://imprintnews.org/featured/uncharted-waters-data-analytics-and-child-protection-in-los-angeles/10867
- Advocacy Child Protective Services Defense, Predictive Analytics in Child Welfare - Helping Hand, or Racial Bias? (Part 2) (2015) https://childprotectiveservicesdefense.com/predictive-analytics-child-welfare-helping-hand-racial-bias-2.html
- Advocacy NCCPR (Richard Wexler), Los Angeles County quietly drops its first child welfare predictive analytics experiment (2017) https://www.nccprblog.org/2017/05/los-angeles-county-quietly-drops-its.html
- Advocacy WitnessLA (Richard Wexler), LA County Nixes Alarmingly Unreliable Predictive Analytics Foster Care Scheme - For Now (2017) https://witnessla.com/op-ed-la-county-nixes-alarming-predictive-analytics-scheme-for-foster-care-for-now/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Escalate checks — State-feedback vigilance
- Require sign-off — Conformity assessment gate
- Gate vendor updates — Vendor quality gate
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Review the riskiest first — Risk-tiered oversight
- Pause AI on alarms — Deployment circuit-breaker
- Upgrade model — Improve the model
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
- Store less data — Data minimization
Documented case histories
- Los Angeles County Project AURA
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model