PAN Lab example
Allegheny Housing Assessment
The score and the scarce bed: a coordinated-entry housing tool
Allegheny County scores people experiencing homelessness on their risk of harm if unhoused, to rank them for housing. Black clients were still served less often.
See more
The Allegheny Housing Assessment (AHA) is risk-scoring software that Allegheny County's Department of Human Services has used since August 2020. From county records, it estimates each client's chance of harms such as a jail booking within a year if they stay unhoused. A score from 1 to 10 then ranks clients for scarce housing.
What the score predicts
The score estimates the chance of four harms within 12 months if the client stays unhoused. They are a mental health inpatient stay, a jail booking, four or more emergency room visits, and another episode of homelessness. The county treats them as signs of harm it would like to prevent.
A separate model estimates each harm. The chances are grouped, added up, and turned into a score from 1 to 10, where 10 means the highest risk. The homelessness outcome was added in August 2025.
The score reads the county's Data Warehouse, which links records from many county agencies. Race is not used. Clients cannot opt out. When a client is assessed again, the new score reads their updated records.
What it replaced
Coordinated entry is the process the U.S. Department of Housing and Urban Development (HUD) requires for ranking people experiencing homelessness for housing. In Allegheny County, the Allegheny Link runs it.
Until August 2020 the Link used the VI-SPDAT (Vulnerability Index Service Prioritization Decision Assistance Tool), an interview survey. The county says it asked intrusive questions about self-harm, drug use, and risky behavior. The county says its answers depended on a person's memory and willingness to share.
The AHA predicts each harm better. The county measures this with AUC, where 50 percent is a coin flip and 100 percent is perfect. For the 2025 version it runs from 66 to 72 percent. The VI-SPDAT ran from 52 to 59 percent.
The self-report questionnaire
Link staff also give the Alt-AHA, a shorter self-report questionnaire, when the warehouse holds less than 90 days of a client's history. The rules name three other cases too. The county estimates about 5 percent of cases lack enough data for a score. The higher of the two scores is the one used.
Who decides
The score does not decide who is housed. It ranks clients on the waiting list. Link staff assess clients and check which programs they are eligible for. They can take a case to case conferencing, where staff discuss an individual client's case, when the score does not fit what a client told them.
The county's homeless resource coordinator matches clients from the top of the list to open program slots. The coordinator is ultimately responsible for deciding who is referred to housing providers. The referral includes the score.
A 2024 study says strict eligibility rules limit how far Link staff can move a client on the list. No rate of changes through case conferencing has been published.
How scarce the housing is
The county assesses far more households than it can serve. In 2019, more than 2,000 eligible households were assessed, about 600 families and 1,450 single adults. About 800 units came open as households left programs. The county could serve fewer than half, a gap of about 1,200 units.
In January 2025 the county had 762 emergency shelter beds, 183 bridge housing beds, 497 rapid rehousing beds, and 1,167 supportive housing beds.
What the 2024 study found
Lingwei Cheng, Cameron Drayton, Alexandra Chouldechova, and Rhema Vaithianathan published a peer-reviewed study in 2024. Vaithianathan led the team that developed the AHA. They studied 6,542 assessments of Black and white single adults from January 2018 to March 2022.
AHA scores were similarly spread across Black and white clients. The racial gap in service rates persisted. A client counted as served when enrolled with a housing program, which does not mean they had moved into housing yet.
White clients were served at 17.6 percent before the AHA and 23.3 percent after. Black clients were served at 14.5 and 19.5 percent.
The study traced the gap at least in part to the self-report questionnaire. White clients were assessed with it twice as often, 24.5 percent against 11.8 percent, largely because the county's data on them was of lower quality. Eligibility factors, such as chronic homelessness, disability, and veteran status, also played a part.
What the 2025 update changed
A team inside the county department updated the AHA in August 2025, adding the homelessness outcome. The update learned from 18,008 assessments made from 2016 through 2024. It raised the AUC for homelessness from 54 to 68 percent.
It also changed who scored highest. Among clients scoring 10, men were 76 percent under the 2025 model and 62 percent under the November 2020 model. Women fell from 34 to 24 percent.
The county says this is likely because men's one-year homelessness risk, 38 percent, is higher than women's, 26 percent. It says more men and fewer women would be assigned housing. It says the racial mix appears similar.
How the tool is checked
Eticas Research and Consulting assessed the model for the county in 2020, before launch. The county says Eticas found it performed better than the VI-SPDAT across all protected groups.
The Homeless Advisory Board oversees the county's Continuum of Care, its HUD-funded network of homeless services. The county consulted that board, clients in shelters, providers, Link staff, and national experts while building the tool.
By 2025 the county checked the scores every day against their usual spread. Scores that look even slightly unusual are not shared until a person checks their quality. The county also shares score data with outside researchers.
What each side says
The county says the AHA is more accurate than the survey and has improved equity in how housing is allocated. People it consulted welcomed dropping questions that made clients revisit traumatic events.
They also worried about clients with little data on record, and that warehouse records may be wrong or incomplete.
The best-known critiques of the county's predictive tools come from the ACLU in 2021 and Virginia Eubanks in 2018. They target the Allegheny Family Screening Tool, the county's child-welfare hotline score, not the AHA. Critics extend their concern about profiling people in poverty from public-service records to the approach the AHA shares.
The county and Eticas stress the gains in accuracy and similar scores across races. The study's authors stress that similar scores did not bring similar service. Both are documented.
What this case shows
The score is computed from records that earlier housing decisions help shape. Later models learn from those records, so today's referrals can shape tomorrow's scores.
A more accurate score, similar across races, did not by itself make service equal. In a decision that rations housing, the gap can come from what surrounds the score. Here that means the questionnaire, the eligibility rules, and the choice of what to predict.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it goes on.
Before any tool is used, six failure pathways are open. The network is at a tipping point, meaning mistakes can build on one another across it. Self-correcting means its mistakes are corrected rather than building on each other.
Lingering effects is a Dynamics setting in which damage outlasts its cause. It is on by default, and always under All Governance Targets. Service and Safety Targets turns it off.
This case has a budget of 9 units. Explore (No Targets) sets no targets. There, the cheapest single tools that make the network self-correcting are Escalate checks, Assign a challenger, and Mark AI-written records, at 2 units each. Store less data and Understand the system each do it alone for 3 units. At their stronger settings, Review on schedule does too, and so does Upgrade model while lingering effects are on.
Under Service Targets Only, each of these also meets the service target, which asks that the score be helping the work. Many other combinations within the budget do too.
Service and Safety Targets and All Governance Targets also ask you to close every failure pathway and refill the Privacy gauge, among other targets. At both levels, the targets can be met one way, using all 9 units. It is Escalate checks, Assign a challenger, Mark AI-written records, and Store less data, each at its standard setting.
Escalate checks closes “Score in the Link data system”. Assign a challenger closes “Assessed clients put on the list”. Mark AI-written records closes “County records used in scoring”, the pathway that drains the Privacy gauge, and “Earlier records read at assessment”. Store less data closes “Assessments recorded in HMIS” and “Referral decisions recorded”.
Understand the system costs 3 units under the two lower levels, 4 under the two higher ones, and 6 at its stronger setting. While it is on, Review the riskiest first, Review on schedule, Mark AI-written records, and Keep skills sharp each cost 1 unit less. At its stronger setting they cost 2 less, but never less than 1 unit. It is in no combination that meets the two higher levels.
While lingering effects are on, Store less data and Review the riskiest first work at reduced strength unless Understand the system is also on. Store less data still closes both its pathways.
More is not better here. Every tool at its strongest setting closes every failure pathway, but costs 29 units, more than three times the budget. It also leaves the score no longer helping the work, so it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the AHA-class coordinated-entry scoring tool network: 6 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 6 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
The county names a second role: the homeless resource coordinator, who matches clients from the top of the priority list and decides who is referred. This example draws that role as its own group of staff. It assumes the coordinator knows a client mostly through the list and the Link staff's assessment. Two roles in sequence raise different questions from one person deciding with a score.
- assumed
This example follows the coordinated-entry scoring documented in the case file on the Allegheny Housing Assessment. It does not rebuild the real tool.
- assumed
This example includes case conferencing, where staff discuss an individual client's case, as the Link staff's way to question a score. A 2024 study found their ability to change a client's place on the waiting list limited by strict eligibility rules. No source publishes how often case conferencing changes a rank. So this example treats it as a limited way to correct a rank.
- assumed
The score reads the county's linked records, which clients cannot opt out of. This example counts that reading as a privacy exposure. When the records hold less than 90 days of history, Link staff also give the Alt-AHA questionnaire. The county estimates about 5 percent of cases lack enough data for a score. The higher score is used. In a 2024 study of single adults, white clients were assessed with it about twice as often as Black clients. The case file records that, and this example does not compute it.
- baseline
This example assumes from the start that earlier housing decisions shape the records the models later learn from. The case file describes this loop. The 2025 update learned from assessments made from 2016 through 2024, years that include the score's use.
- assumed
This example places the priority list between the score and the Link staff, because the tool's output is a ranked list. It has no pathways of its own, so it changes nothing in how mistakes move.
- assumed
The case file records the racial gap in service rates, the uneven use of the Alt-AHA by race, and the 2025 shift toward men among the top scores. This example shows how mistakes move among the score, the staff, and the records. It does not model groups of people or estimate unequal harm to the people experiencing homelessness whom the score ranks.
What this example does not show
Show all 2 limitations
- This example shows how mistakes move through records, ranking, and the staff who use them. It does not model groups of people or estimate unequal harm to the people experiencing homelessness whom the score ranks. The case file records the racial gap in service rates, the uneven use of the Alt-AHA by race, and the 2025 shift toward men among the top scores. They were measured outside any network like this one.
- A housing referral is a one-time decision about a scarce unit. This example follows mistakes passing between parts. So it shows two things: the deployment's checks, such as the external assessment and daily score monitoring, and the loop from records back to the score. It makes no claim about any one person's outcome.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A peer-reviewed 2024 evaluation of the Allegheny Housing Assessment found that although the tool was substantially more accurate than the VI-SPDAT survey it replaced and produced similar risk-score distributions across race, it did not reduce the racial disparity in service rates: white single adults were served at about 23.3% versus 19.5% for Black clients.
empirical- Academic Cheng, Drayton, Chouldechova and Vaithianathan, Algorithm-Assisted Decision Making and Racial Disparities in Housing: A Study of the Allegheny Housing Assessment Tool (Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society; arXiv:2407.21209) https://arxiv.org/abs/2407.21209
- Government Allegheny County Department of Human Services (Allegheny Analytics), Allegheny Housing Assessment (AHA) Frequently Asked Questions (January 2026) https://analytics.alleghenycounty.us/wp-content/uploads/2026/01/AHA-FAQs-Update_Jan_2026.pdf
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Housing & homelessness services domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Review on schedule — Oversight cadence & retrospectives
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Assign a challenger — Structured dissent
- Mark AI-written records — Provenance labeling
- Store less data — Data minimization
- Upgrade model — Improve the model
- Understand the system — Understand the system
Documented case histories
- Allegheny Housing Assessment
- VI-SPDAT
- LA's coordinated-entry triage revision: the fix that needed fixing
- LA County Homelessness Prevention Unit
- Santa Clara County Homelessness Prevention System
- Homebase Risk Assessment Questionnaire
- Xantura OneView (predictive homelessness flagging)
- London's Strategic Insights Tool: one linked memory of rough sleeping read by every borough
- CHAI (chronic-homelessness prediction)
- Calgary Drop-In Centre: interpretable screening a shelter's own staff choose to check
- San Jose's camera car: a low-precision detector aimed at who is sleeping outside
- Imagine LA Benefit Navigator copilot
- SafeRent Tenant Screening Score
- CrimSAFE criminal-record tenant screening
- One engine, many rivals: a shared rent-setting model and the record it writes back