PAN Lab example
LA's coordinated-entry triage revision
Two scores in one queue: the transition that let a retired bias back in
Los Angeles swapped a biased housing survey for a fairer score. LAHSA says clients qualified more easily on the old one, so providers kept it.
See more
The Los Angeles Homeless Services Authority (LAHSA) uses a survey to score single adults experiencing homelessness for housing. In 2024 it began replacing the Vulnerability Index-Service Prioritization Decision Assistance Tool (VI-SPDAT), found to under-score clients of color, with the Los Angeles Housing Assessment Tool (LA HAT). Until mid-2026 both ran, and LAHSA says the older one qualified clients more easily.
How the new score was built
The CESTTRR project, short for Coordinated Entry System Triage Tool Research and Refinement, built the LA HAT between March 2020 and May 2023. USC's Center for AI in Society led it with the California Policy Lab at UCLA. LAHSA's Ad Hoc Committee on Black People Experiencing Homelessness prompted the work. A Community Advisory Board and a Core Planning Group governed its design choices.
Each question's point weight came from a statistical model fitted to past data. The December 2024 press account calls it ordinary least squares regression. The project's technical appendix describes a lightly penalized Lasso model with positive weights.
The model learned from 71,747 past VI-SPDAT assessments of 54,543 people, taken from July 2015 to October 2018. They were linked to records from seven county agencies. The outcome was a two-year measure of serious harm, defined with the community. It counted an emergency or inpatient visit, a crisis-stabilization episode, justice involvement, a substance-use-disorder diagnosis, or death.
The questions were cut from 35 to 19. The algorithm picked 9, and the community boards put back about 10. The report says putting them back did not meaningfully change accuracy or fairness. The final report runs to 133 pages.
How the survey is given
A trained case manager gives the survey face to face and records the client's own answers. It can be done on paper and keyed later into the Homeless Management Information System (HMIS), the shared homeless-services database. Some records go into a comparable database for domestic-violence services instead.
The score uses only those answers. County records were used to set the weights, never to score a person live.
The report recommends trauma-informed practice: not at intake, read word for word, and in a private place. In a 2022 to 2023 pilot with 9 agencies, 18 case managers, and 49 clients, both clients and case managers preferred the revised tool.
A separate USC-built algorithm for matching people to housing was deliberately not deployed. Los Angeles swapped only the scoring survey and kept human discretion over matching.
Accuracy traded for fairness
The CESTTRR research found the VI-SPDAT barely beat a coin flip at spotting who would come to serious harm. Its area under the curve, a score where 0.5 is chance and 1.0 is perfect, was 0.54.
It also missed more real cases among clients of color. Its generalized false-negative rate, the share of real cases it missed, was 54 percent for white participants. It was up to 8.5 percentage points higher for Black, Latinx, and Native Hawaiian or Pacific Islander participants. So vulnerable clients of color were systematically scored too low.
The LA HAT was deliberately made less accurate to make it fairer. A version tuned only for accuracy reached 0.64. The equity-adjusted version put into use scores 0.60. It narrows the gap in missed cases from 5.9 to 0.7 points for Black clients, and from 3.2 to 0.2 points for Latinx clients. Of the clients it scored in the top tenth, 57 percent went on to serious harm, against 49 percent under the VI-SPDAT.
The report is candid about its limits. A hypothetical model with far richer data could reach 0.831. The report puts the best possible generalized false-negative rate at 42 percent.
Every one of these figures is a pre-deployment estimate on 2015 to 2018 data set aside for testing. None is a result observed after launch.
The changeover
CES is Los Angeles's Coordinated Entry System, the largest homeless-services system in the United States. In 2024 its Policy Council approved the LA HAT for permanent supportive housing consideration alongside the VI-SPDAT. Existing VI-SPDAT scores stayed valid, to lessen the disruption of the change.
LAHSA fully launched the tool in 2025. It ran 94 training sessions with its partner A.C.T.I.O.N. to Healing and trained 3,139 staff, past a goal of 3,000. It opened access across 688 eligible HMIS programs and translated the survey into the nine languages the county serves by rule.
For a time the two surveys ran side by side with different bars. A client qualified for consideration at 8 or more on the VI-SPDAT, or 17 or more on the LA HAT. LAHSA's own page dates this two-threshold period from December 2025 to April 2026.
By LAHSA's account, initial quantitative data and provider feedback showed clients were more likely to score eligible on the VI-SPDAT. So direct-service providers opted to give the VI-SPDAT instead. LAHSA says this trend "perpetuated the racial bias of the VI-SPDAT in the System" and slowed the new tool's uptake. LAHSA has not released the numbers behind that account.
How the council corrected it
The fix was a change of rules, not a better model or new training. On April 22, 2026 the CES Policy Council lowered the LA HAT threshold from 17 or more to 12 or more. Where a client has both scores, the most recent LA HAT score now governs. Programs with LA HAT access had to stop new VI-SPDAT surveys.
The VI-SPDAT was switched off for those programs on May 1, 2026. It was removed for new surveys system-wide on June 30, 2026. The council will decide in fall 2026 when older VI-SPDAT scores stop counting.
The current matching guidance was revised on May 27, 2026. It says plainly that "the VI-SPDAT is being phased out of the System due to its inherent bias." It also says clients with both scores who were eligible on April 22, 2026 are not made ineligible by the change.
Only a client's most recent assessment counts. The sources document no way for a case manager to override a computed score, so the threshold rule applies mechanically. Human judgment works outside the score. Case conferencing is another route to housing readiness, and caseworkers keep discretion over matching.
Who watches the results
LAHSA and the CES Policy Council monitor assessment and prioritization outcomes. The matching guidance names partners in that work: A.C.T.I.O.N. to Healing, Arc4Justice, Papalotl Consulting, and paid lived-experience bodies.
Arc4Justice is commissioned to publish a first evaluation of the tool's rollout in 2026. No evaluation of outcomes after launch has been published.
The shortage behind the list
The score decides who is considered for a scarce benefit. Late 2024 reporting found that for every available slot in permanent supportive housing in Los Angeles County, about four more are needed. About 75,000 people were unhoused, up from about 53,000 in 2018. About 17,000 were waiting for permanent supportive housing. Black people were under 10 percent of the county's population but over 30 percent of its unhoused people.
What each side says
Community members with lived experience shaped the effort and also criticized it. Advisory board member Reba Stevens argued that "everybody is vulnerable" and that scoring cannot fix the housing shortage. Lead researcher Eric Rice acknowledged that he had helped make "a system that is inadequate... fair, or more fair," not a solved one.
What this network is drawn from
This network follows the pattern the case file describes. It does not copy the actual tool. It draws the changeover: the two surveys, the case managers who give them, the provider agencies, the council, and the HMIS record. It also draws two checks the sources do not describe during the changeover. One compares the two surveys, and one reviews which survey providers gave. The people scored stay outside the network.
The LA HAT replaced the VI-SPDAT in the single-adult system. Family and youth systems use other tools. The Lab's separate VI-SPDAT case covers the older survey on its own, across many US communities.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it may go on.
This case has a budget of 11 units and starts with no pressure applied. Explore (No Targets) sets no targets. The other three levels all ask for the network's mistakes to be contained, meaning corrected rather than building on each other.
Under Service Targets Only, the targets can be met in hundreds of ways. The cheapest is Escalate checks alone, for 2 units. It is the only tool that meets them on its own. Many pairs cost 4 units: Escalate checks with Mark AI-written records, Peer sharing rules, Review on schedule, or Keep prompts neutral. Some ways leave Escalate checks out, such as Mark AI-written records with Gate record entries for 5 units.
Service and Safety Targets also asks you to close every failure pathway, keep up with the work, and refill the Privacy gauge. Six pathways are open before any tool is used. They are LA HAT score sets eligibility, VI-SPDAT score sets eligibility, Answers entered into HMIS, Stored scores order referrals, Case manager picks the survey, and Agency survey practice.
The targets can be met in two ways, each costing the whole budget of 11 units. Both use Escalate checks, Mark AI-written records, Peer sharing rules, and Keep prompts neutral. The fifth tool is Store less data or Gate record entries.
Escalate checks closes the two links from the scores to case managers. Mark AI-written records closes Stored scores order referrals. Peer sharing rules closes Agency survey practice. Keep prompts neutral closes Case manager picks the survey. Store less data or Gate record entries closes Answers entered into HMIS. No stronger setting fits within the budget.
All Governance Targets asks for more again. The same two tool sets meet it, at 11 units each.
Check with a second model, Review on schedule, and Upgrade model are in no combination that meets the two higher levels within the budget. With the budget lifted, each can join the five tools, for 13 or 14 units. Pause AI on alarms is in no combination that meets any level, even with the budget lifted. It cuts the benefit the scores bring too far.
More is not better here. Using every tool at its strongest setting costs 39 units, well over the budget of 11. It contains the mistakes and closes every failure pathway. But the benefit the scores bring falls below what the targets ask, so it meets them at none of the three levels.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the LA-HAT-class coordinated-entry triage-revision system network: 6 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This network follows the changeover pattern the LAHSA case file describes. It is not a copy of the actual tool. It draws two scoring surveys running side by side: the corrected LA HAT and the older Vulnerability Index-Service Prioritization Decision Assistance Tool (VI-SPDAT). A frontline case manager chooses which one to give. It covers the changeover only. The Lab's separate VI-SPDAT case covers the older survey on its own. A USC-built housing-matching algorithm, which Los Angeles chose not to deploy, is also left out. What drives this network is the case manager's choice between two surveys, not either score's accuracy.
- baseline
The network assumes the two surveys qualified clients unevenly while both ran. A client qualified for supportive housing consideration at 8 or more on the VI-SPDAT, or 17 or more on the LA HAT. LAHSA says initial quantitative data and provider feedback showed clients were more likely to score eligible on the VI-SPDAT. So direct-service providers opted to give it instead. LAHSA says this trend perpetuated the VI-SPDAT's racial bias in the system and slowed the new tool's uptake. The fix was a change of rules, not a better model or new training. On April 22, 2026 the CES Policy Council lowered the LA HAT threshold to 12 or more. It ruled that the most recent LA HAT score supersedes a VI-SPDAT score where both exist. It ended new VI-SPDAT surveys, for programs with LA HAT access on May 1, 2026 and system-wide on June 30, 2026. The council will decide on phasing out older VI-SPDAT scores in fall 2026.
- baseline
The network assumes two checks were missing while both surveys ran. First, nothing compared how many clients qualified on each survey. So the easier survey took over decisions unseen, until quantitative data showed it. Second, nothing reviewed which survey providers gave, or to whom. So the pattern ran until it was found, rather than being caught early. The redesign itself was careful. It had a 133-page technical report, a negotiated trade of accuracy for fairness, questions the community put back, and trauma-informed administration. The changeover failed where no one was watching: the gap in how easily each survey qualified people.
- baseline
The corrected survey was deliberately made less accurate to make it fairer, and its report is candid about its limits. The LA HAT's area under the curve, a score where 0.5 is chance, is 0.60, down from 0.64 for an accuracy-only version. It narrows the racial gap in missed cases from 5.9 to 0.7 percentage points for Black clients, and from 3.2 to 0.2 for Latinx clients. These are pre-deployment estimates on 2015 to 2018 data set aside for testing. The report places it well below the 0.831 a hypothetical model with far richer data could reach. The report puts the best possible generalized false-negative rate at 42 percent. So this is not a story about a bad model. The successor was more careful and fairer than the survey it replaced, and it still did not take hold. The bias came back through the changeover, not the model. That is why Upgrade model is the wrong lever here.
- assumed
The network assumes the privacy exposure is the self-report intake. A case manager gives the survey and keys the answers into the Homeless Management Information System (HMIS), or a comparable domestic-violence database. The survey uses no live county data. County records were used only to set its weights, so the network draws no link from the record back into the scores. The client's own answers cover housing history, health, and justice involvement. They are entered into a shared record from which the prioritization list is built. Store less data and Gate record entries act on that entry. No score denies anyone automatically, and the sources document no override of a computed score. The threshold rule applies mechanically. Human judgment works outside the score, through case conferencing as another route and discretion over matching.
- assumed
The people experiencing homelessness whom the surveys score are not in the network. It follows how mistakes pass between the surveys, staff, and records, not any person's housing outcome. The harms it shows are institutional: the biased survey taking over decisions during the changeover, an old score shaping the list, and an unwatched gap between the surveys. The racial harm is documented outside the network. Every fairness and accuracy figure is a pre-deployment estimate on 2015 to 2018 data. The LA HAT's figures are its builders' predictions, not results observed after launch. No evaluation of outcomes after launch has been published, though Arc4Justice is commissioned to publish a first-phase implementation evaluation in 2026. LAHSA's account of the uneven eligibility cites initial quantitative data without releasing the numbers. A score, rank, or threshold here is an institutional signal, never a person.
What this example does not show
Show all 4 limitations
- This example does not show any person's housing outcome. It shows how mistakes pass between the surveys, the staff who give them, and the HMIS record. The people scored stay outside it. The harms it shows are institutional: the biased survey taking over decisions during the changeover, an old score on the list, and an unwatched gap between the surveys. The racial harm the case is about is documented outside this example.
- This example does not show how the surveys performed after launch. Every fairness and accuracy figure comes from the 2023 final report of the Coordinated Entry System Triage Tool Research and Refinement (CESTTRR) project. Each is a pre-deployment estimate on 2015 to 2018 data set aside for testing. The VI-SPDAT's area under the curve was 0.54, near chance, with gaps in missed cases of up to 8.5 percentage points. The LA HAT's was 0.60, traded down from 0.64 for fairness. Its predicted gaps fall from 5.9 to 0.7 points for Black clients and from 3.2 to 0.2 for Latinx clients. No evaluation of outcomes after launch has been published, though a first-phase implementation evaluation is commissioned for 2026. How this example looks when it opens is not a safety finding about any real deployment.
- This example does not show how large the eligibility gap was. LAHSA's central claim is its own statement. It says the two thresholds made clients more likely to qualify on the VI-SPDAT, so providers gave it, and this perpetuated the survey's racial bias. It cites initial quantitative data and provider feedback but releases no numbers. The direction is documented, but the size is not. The rollout is best described as phased from 2024 to 2026, not tied to one start date. LAHSA's own page dates the two-threshold period from December 2025 to April 2026.
- This example covers the changeover, not the older survey on its own. The Lab's separate VI-SPDAT case covers that survey across many US communities, and this example refers to it without repeating it. It also leaves out a separate housing-matching algorithm built at the University of Southern California (USC). Los Angeles deliberately did not deploy it, swapped only the scoring survey, and kept human discretion over matching. The deployed survey uses the client's own answers and no live county data. This example is about the two-survey changeover and the choice it created, not about how either model works inside.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The Los Angeles Coordinated Entry System replaced the VI-SPDAT survey for single adults with the Los Angeles Housing Assessment Tool, a 19-item self-report score whose weights were derived by a regression on 71,747 historical assessments linked to county records; where the CESTTRR research estimated the VI-SPDAT scored near chance (AUC 0.54) with racial false-negative gaps up to 8.5 percentage points, the equity-adjusted successor was deliberately traded down in overall accuracy (AUC 0.60, from an accuracy-only 0.64) to close those gaps to under one percentage point, and every such figure is a pre-deployment estimate on 2015 to 2018 held-out data rather than an observed post-launch outcome.
empirical- Academic Rice, Milburn, Vayanos, Rountree, Hill, Petering, Blackwell, Santillano and colleagues, CESTTRR Coordinated Entry System Triage Tool Research and Refinement Final Report (USC Center for Artificial Intelligence in Society, 2023) https://cais.usc.edu/wp-content/uploads/2023/11/CESTTRR-Final-Report-2023.pdf
- Government Los Angeles Homeless Services Authority, Los Angeles Housing Assessment Tool (LA HAT) (2025) https://www.lahsa.org/news?article=1033-los-angeles-housing-assessment-tool-la-hat-
During the dual-tool transition the two instruments' PSH-consideration thresholds were 8-plus on the VI-SPDAT and 17-plus on the LA HAT, and by LAHSA's account initial quantitative data and provider feedback showed participants were more likely to obtain an eligible score under the VI-SPDAT, so direct-service providers opted to administer it, a trend LAHSA states 'perpetuated the racial bias of the VI-SPDAT in the System'; on April 22, 2026 the CES Policy Council lowered the LA HAT threshold to 12-plus, ruled the most recent LA HAT score supersedes a coexisting VI-SPDAT score, and forced deactivation of new VI-SPDAT completions (for LA HAT-access programs on May 1, 2026 and system-wide on June 30, 2026), though LAHSA has not released the underlying eligibility-rate numbers.
empirical- Government Los Angeles Homeless Services Authority, Los Angeles Housing Assessment Tool (LA HAT) (2025) https://www.lahsa.org/news?article=1033-los-angeles-housing-assessment-tool-la-hat-
- Government Los Angeles Homeless Services Authority, Los Angeles Housing Assessment Tool (LA HAT) Spring 2026 Implementation Updates (2026) https://www.lahsa.org/documents?id=9877-los-angeles-housing-assessment-tool-la-hat-implementation-improvements-spring-2026-
- Government Los Angeles Homeless Services Authority and the LA CES Policy Council, CES Permanent Supportive Housing Prioritization and Matching Guidance (2026) https://www.lahsa.org/documents?id=7658-ces-psh-prioritization-and-matching-guidance-effective-07-01-2026-.pdf
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Housing & homelessness services domain page.
Levers available here and the patterns behind them
- Check with a second model — Cross-model verification
- Pause AI on alarms — Deployment circuit-breaker
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Peer sharing rules — Peer-edge governance
- Escalate checks — State-feedback vigilance
- Keep prompts neutral — Framing and mirroring reduction
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- LA's coordinated-entry triage revision: the fix that needed fixing
- Allegheny Housing Assessment
- VI-SPDAT
- LA County Homelessness Prevention Unit
- Santa Clara County Homelessness Prevention System
- Homebase Risk Assessment Questionnaire
- Xantura OneView (predictive homelessness flagging)
- London's Strategic Insights Tool: one linked memory of rough sleeping read by every borough
- CHAI (chronic-homelessness prediction)
- Calgary Drop-In Centre: interpretable screening a shelter's own staff choose to check
- San Jose's camera car: a low-precision detector aimed at who is sleeping outside
- Imagine LA Benefit Navigator copilot
- SafeRent Tenant Screening Score
- CrimSAFE criminal-record tenant screening
- One engine, many rivals: a shared rent-setting model and the record it writes back