PAN Lab example
Sistema Alerta Niñez (Chile)
Scored before anyone knocks: a child-risk targeting tool
Chile's Sistema Alerta Niñez ranks children by risk, from data families gave for benefits. Local offices use it to decide who gets early help first.
See more
Sistema Alerta Niñez is a predictive risk tool built for Chile's Ministry of Social Development and Family. It estimates each child's risk of a future violation of their rights from 280 administrative variables, and ranks children by that risk. Local Childhood Offices use the ranking to decide whom to offer early help first.
Who built it
The tool grew out of President Piñera's 2018 national agreement on childhood, the Gran Acuerdo Nacional por la Infancia. A consortium of two university groups built the model in 2018 and 2019. They were GobLab at Universidad Adolfo Ibáñez and the Centre for Social Data Analytics at Auckland University of Technology.
The Chilean firm Actis implemented and maintained the model. The developers explored several methods and chose LASSO regression. It is a standard statistical method that keeps only the variables that help the prediction.
How well it predicted
In a 2019 proof of concept, the models reached an area under the curve, or AUC, of roughly 0.88 to 0.95 on test data. AUC measures how well a model ranks cases that had the outcome above cases that did not. A score of 1 is perfect, and 0.5 is no better than chance.
The outcome was a child's separation from family, or contact with child-protection programs, within two years. The deployed model's real-world performance was never publicly released.
How it is used
The Local Childhood Offices, Oficinas Locales de Niñez or OLN, use the prioritized list to decide whom to offer early help first. Officially the score is one more input, un insumo más, always subordinate to the judgment of OLN professionals.
Sistema Alerta Niñez is also the platform where OLN staff register and monitor cases and interventions. The model does not use the field knowledge they gather there.
The tool was piloted in 12 communes, Chile's municipalities. About 3,354 children had been served there by August 2020. The network of offices later expanded toward a majority of the country's communes.
Where the data comes from
Families supplied the data to qualify for social benefits, through the household registry, the Registro Social de Hogares. The 280 variables come from benefit, education, health, child-protection, crime, and census records.
The predictive use was recast as ordinary targeting of services, focalización in Spanish. So the people scored were never told a risk ranking existed. They could not opt out, and they were not consulted.
What was never made public
An outside audit of the ranking for algorithmic bias was carried out in 2020. The case file reports that the Inter-American Development Bank funded it and that the consultancy Eticas conducted it. Its criteria and results were never made public.
The sources also list evaluations by the World Bank and the United Nations Development Programme. They do not say what those evaluations found.
What the developers acknowledged
The developers acknowledged that the model is less able to identify higher-income children at risk. The reason they gave is that lower-income families have more contact with the state. This is their own qualitative account, not a measured disparity.
Why this case matters
Most cases in the case files turn on the output side: who acts on a score and who can correct it. This one turns on the input side, where data given for benefits becomes a risk ranking nobody was told about.
The case file names controls that might have bound the system but were never turned on. The independent bias audit was commissioned and then withheld. The model's operational performance was never published. The model never used the field knowledge OLN staff gather.
The case file draws the lesson plainly. A score whose accuracy is never disclosed cannot be contested. Transparency at the input and evaluation boundaries, not accuracy at the output, is where a tool like this is governed.
Where the facts come from
The facts here come from the case file and the sources behind it. They include two publications by Matias Valderrama for Derechos Digitales: a 2021 report in Spanish and a 2022 report in English. They also include the Centre for Social Data Analytics' 2019 proof of concept.
Victoria Adelmant wrote about the system in 2022 for the Center for Human Rights and Global Justice at NYU School of Law. Paz Peña wrote a 2022 case study of it for Not My AI.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.
This case has a budget of 10 units. Each tool costs the same at every target level.
Explore (No Targets) sets no targets. Before any tool is used, the network keeps its mistakes contained: they are corrected rather than building on each other. Under Service Targets Only, the targets are met before any tool is used. That level asks for mistakes to be contained and for the ranking to be helping the work.
Under Service and Safety Targets and All Governance Targets, the targets can be met. Both levels also ask you to close every failure pathway and to bring privacy protection back to its target. Before any tool is used, three pathways are open: “280 variables scored”, “Ranking to OLN staff”, and “Cases logged to platform”. Only the first drains privacy protection.
The cheapest combination costs 7 of the 10 units, with each tool at its standard setting. Escalate checks has OLN staff check the ranking more closely when monitoring flags trouble, which closes “Ranking to OLN staff”. The sources do not describe such monitoring here. Store less data closes “Cases logged to platform”.
Mark AI-written records marks machine-made content in records so that their readers, the model included, can weigh it. In the Lab it closes “280 variables scored”, because it acts on every stored record the model reads. The sources do not describe any of those 280 variables as written by a machine. So this result comes from how the Lab draws the tool, not from the deployment.
Under Service and Safety Targets, every combination that meets the targets includes Escalate checks and Store less data. Vet connections can take the place of Mark AI-written records there, for 8 units. Of the 2,392 combinations of tools and settings within the budget, 37 meet these targets. They use 19 different sets of tools.
Lingering effects is a Dynamics setting in which damage outlasts its cause. Under All Governance Targets it is always on, and then Vet connections changes nothing here. Every combination that meets these targets includes all three tools of the cheapest combination. Of the same 2,392 combinations, 29 meet them, using 13 different sets of tools. That level also asks for the ranking to be clearly helping the work.
More is not better here. Every tool at once, each at its stronger setting where it has one, costs 40 units, four times the budget. It closes every failure pathway but leaves the ranking helping the work too little, so it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the SAN-class predictive child-risk targeting tool network: 6 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This network follows the predictive-targeting pattern documented in the Sistema Alerta Niñez case file. It does not rebuild the actual tool.
- baseline
The network draws the cross-agency child data as one of its parts: the 280 administrative variables the model scores from. Their pathway into the model is marked as sensitive for privacy. Families supplied that data to access benefits, without informed consent to the risk ranking or a way to opt out.
- assumed
The network draws a pathway from the SAN case platform back into the model. The sources record that the model does not use OLN's field knowledge, so the deployment never built this loop. Built without care, it would let the model learn from children its own ranking chose for outreach.
- assumed
The network draws the independent bias audit as a reviewer, with a check on the OLN staff's use of the ranking. The sources record that the audit was carried out, but its criteria and results were never made public. The case file reads it this way: an audit whose results no one may see is not a control.
- assumed
The priority ranking list is drawn on the pathway from the model to the OLN staff. The case file describes a system whose output was a prioritized list. In the Lab the list is the label on that pathway. Officially the score is one more input, always subordinate to OLN professionals' judgment.
- assumed
The developers acknowledged a socioeconomic gradient: the model is less able to identify higher-income children at risk. That is their own qualitative account, not a measured disparity published by an independent audit. The network traces how mistakes pass between the model, the staff, and the records, not between groups of people. It estimates no unequal harm to the children served.
What this example does not show
Show all 1 limitation
- The documented concern is a socioeconomic gradient. The developers acknowledged the model is less able to identify higher-income children at risk, because lower-income families have more contact with the state. That is their own account, not a measured disparity published by an independent audit. The Lab traces how mistakes pass between parts of the deployment, not between groups of people. It estimates no unequal harm to the children served. The concern is recorded in the sources behind the case file, and any measurement of it lies outside a network like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Sistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits, without informed consent to the risk ranking or a way to opt out; the model's developers acknowledged it was less able to identify higher-income children at risk, because lower-income families have more contact with the state.
empirical- Investigative Derechos Digitales (Matias Valderrama), IA e inclusion: Chile 'Sistema Alerta Ninez' y la prediccion del riesgo de vulneracion de derechos de la infancia (2021) https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf
- Investigative Derechos Digitales (Matias Valderrama), AI and Inclusion: Chile 'The Child Alert System' (2022) https://www.derechosdigitales.org/wp-content/uploads/02_Informe-Chile-EN_180222.pdf
- Academic Center for Human Rights and Global Justice, NYU School of Law (Victoria Adelmant), Risk Scoring Children in Chile (2022) https://chrgj.org/2022-04-20-risk-scoring-children-in-chile/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Escalate checks — State-feedback vigilance
- Vet connections — Connection authorization
- Store less data — Data minimization
- Gate vendor updates — Vendor quality gate
- Assign a challenger — Structured dissent
- Review on schedule — Oversight cadence & retrospectives
- Require sign-off — Conformity assessment gate
- Keep skills sharp — Deskilling-arrest mandate
- Review the riskiest first — Risk-tiered oversight
- Mark AI-written records — Provenance labeling
- Upgrade model — Improve the model
Documented case histories
- Sistema Alerta Niñez (Chile)
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model