PAN Lab example
London's Strategic Insights Tool
One shared memory and thirty-three readers: consolidating a city's rough-sleeping records
London's Strategic Insights Tool links three record systems into one picture of people sleeping rough. All 33 local authorities read it to plan services.
See more
The Strategic Insights Tool for Rough Sleeping matches each person's records from three London systems into one history of their contacts with services. Faculty built the matcher. Homeless Link has hosted and maintained it for the Greater London Authority since February 2024. The project reports that it misses about 9 in 100 true matches, so some counts run low.
The three sources
CHAIN is London's street-outreach database. The Greater London Authority commissions it, Homeless Link runs it, and an automated pipeline sends its records to the tool every week. In-Form holds charity accommodation and hostel casework, which service providers send every month by upload or an automated link. H-CLIC holds borough statutory homelessness applications, which boroughs send every three months, with history from 1 January 2022.
Each source has its own custodian and its own update schedule. The tool links them on fuzzy-matched names, National Insurance numbers, dates of birth, and phone numbers.
How the matching works
The matcher joins two records only when it rates them at least 85 percent likely to belong to the same person. That keeps false positives, two different people's records matched by mistake, as few as possible. When linked records disagree, it takes the detail from the most reliable source.
The project's own second-phase data protection impact assessment, version 2.0, dated 23 October 2023, reports 91 percent recall. A data protection impact assessment is the review of privacy risks that UK data protection law requires for high-risk processing. That means it finds 91 of every 100 true matches. It was published on the LOTI site in April 2025. LOTI is the London Office of Technology and Innovation.
The assessment says: "we miss 9/100 matches and numbers subsequently appear lower in places where they should be higher." It adds that recall will vary as new data of varying quality comes in. No false-positive rate is published. The delivery team reports these figures itself.
What the tool is not
It makes no decision about any person and runs no risk score. Users see aggregate views of whole populations, plus records their own organisation uploaded. The project states it "is not a substitute for published data and reports." Its purpose is to inform commissioning, strategy, and funding for groups of people.
LOTI says users are told the numbers may differ from their expectations, because of data quality and the matching threshold.
How it was built
The Life off the Streets Partnership of the Greater London Authority and London Councils started the project in 2022, with advice from Bloomberg Associates. The applied-AI firm Faculty became technical delivery partner and sole data processor on 5 June 2023. The homelessness social enterprise Beam supported delivery.
A first version went live on 8 September 2023, piloted with Camden, Hillingdon, Lambeth, and Westminster and their service providers. An information-governance review of that pilot, completed 31 October 2023, cleared the wider rollout. The tool reached all London boroughs and more service providers by 29 February 2024.
Faculty's contract ended on 2 February 2024. The GLA then contracted Homeless Link, which already runs CHAIN, to host, manage, and maintain the tool. By March 2025 LOTI reported 45 contributing organisations, and the tool runs on Amazon Web Services.
Who is responsible
Every participating borough and charity, the GLA, and London Councils are joint data controllers under a pan-London Data Sharing Agreement. Each shares legal responsibility for how the linked data is used. The data protection impact assessment sets out each organisation's position. LOTI's pan-London information-governance lead wrote data-protection guidance for charities and other providers.
How privacy is handled
The people whose records are linked are not notified individually. The legal grounds the partners cite are not consent. They are a public task or a legitimate interest. For sensitive data, they are substantial public interest and research.
In user testing, a user stacked filters and got one view down to 1 or 2 people. Following Office for National Statistics practice, any output of 5 or fewer now shows as "equal to or less than 5".
Records are kept five years per person, in line with how the government's rough-sleeping indicators are defined. Someone not seen sleeping rough for five years is treated as new if they return. Unmatched service-provider records are deleted automatically.
A first-phase review found several sensitive fields unused. They include substance misuse, current mental health concerns, pregnancy status, prison history, care-leaver history, and entitlement to welfare benefits. The project kept collecting them for planned features, and they are "not released or visible to users in any way." What it removed was data outside the request, and the unmatched records.
What the evidence shows
Faculty's case study says 40 organisations have supplied data, with 151 onboarded users. Its headline figures are 32 boroughs onboarded and 12 service providers supplying data. It says the work gave LOTI "an understanding of the rough sleeper population for the first time," and replaced manual data tasks. These are vendor claims.
The Chief Digital Officer for London's first-year review, December 2024, says the tool helped commissioners "challenge assumptions about local rough sleeping patterns." It says the tool revealed "previously hidden connections between street homelessness and Housing Options services." It names the funders as London Housing Directors, the GLA, and the Ministry of Housing, Communities and Local Government. No published figures or independent evaluation located back these claims.
How many people it concerns
Mayoral decision MD3161, 23 August 2023, approved £144,535 of London Councils funding for CHAIN staffing to support the tool, within a £1,401,011 package. It recorded 10,053 people seen sleeping rough in London in 2022-23, up 21 percent on the year before.
Mayoral decision MD3331, 4 February 2025, approved £736,000 to expand CHAIN through Homeless Link for 2025-26 to 2027-28. It recorded 11,993 people seen sleeping rough in 2023-24, up 19 percent. That decision does not name the tool.
What the case asks
Most housing systems in this Lab score a person. This one scores no one. It links separate records into one shared picture that every borough reads. So the matcher's missed matches are not one team's local error. They lower the same figures for the whole city at once, and no second, differently sourced view exists to disagree.
The case asks whether anyone checks the linked journeys against their sources, and whether readers treat the counts as the low estimates they are. A sharper matcher alone would not answer either question.
What the sources do not show
No source gives a false-positive rate, an independent audit of the matching, or any routine reconciliation of the linked journeys against their sources. No source describes an independent evaluation of whether the shared figures improved commissioning.
Forecasting demand, and adding probation, health and care, or eviction data, are stated ambitions in every source. None is a working feature.
Where the facts come from
The facts come from the project's second-phase data protection impact assessment, LOTI's project page, and three LOTI blog posts, one co-written with Faculty. They also come from a Faculty blog on the site of the trade body techUK, Faculty's case study, and the Chief Digital Officer for London's first-year review. Two mayoral decisions, MD3161 and MD3331, give funding and population figures.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.
This case has a budget of 11 units. The network opens in cascading failure, where mistakes spread faster than they are corrected. Contained means self-correcting: mistakes are caught faster than they spread.
Each tool has a standard and a stronger setting. The Lab opens with side-effects on, where governance carries its own costs. It also opens with lingering effects on, where damage outlasts its cause.
Under Explore (No Targets), which sets no targets, no single tool contains the mistakes. The cheapest pairs that do cost 4 units each. They are Escalate checks with Peer sharing rules, Mark AI-written records with Escalate checks, and Mark AI-written records with Keep skills sharp. With lingering effects off, Mark AI-written records with Peer sharing rules also does it for 4 units. Many other pairs contain the mistakes for more.
Under Service Targets Only, the same cheapest pairs also meet the targets. That level also asks for the automated system to be helping the work. Many other combinations meet them too.
Under Service and Safety Targets and All Governance Targets, you must also close every failure pathway and keep the work from being strained, among other targets. Counting stronger settings, two combinations meet the targets under Service and Safety Targets. One meets them under All Governance Targets. Each costs all 11 units.
All of them combine Mark AI-written records, Peer sharing rules, Escalate checks, and Keep prompts neutral. The fifth tool is Gate record entries. Under Service and Safety Targets, Store less data can take its place.
All Governance Targets keeps lingering effects on. There, Store less data at its standard setting leaves the matcher's writes open, and its stronger setting costs too much. Every combination leaves the automated system clearly helping and the work ahead.
Each of the first four closes pathways no other tool here closes. Mark AI-written records closes CHAIN's records to the matcher and the borough teams' reads of stored journeys. Peer sharing rules closes the coordination between boroughs and the pan-London bodies.
Escalate checks closes the tool's views to borough teams and the pan-London bodies. Keep prompts neutral closes the data set and field mapping that shape the matcher. Gate record entries and Store less data each close the matcher's writes and the entries into In-Form and H-CLIC.
The safeguard the case file's reading puts first, reconciling the linked copy against its sources on a rhythm, is in none of these combinations. Check copied records adds the reconciliation check but closes no pathway. Review on schedule, Keep skills sharp, and Upgrade model close no pathway on this network either.
The reading's other two safeguards are in all of them. Mark AI-written records keeps the undercount on every read, and Peer sharing rules adds the independent evaluation.
More is not better here. Every tool at its strongest setting, ignoring the budget, closes every pathway. It still misses All Governance Targets, because the automated system is then not clearly helping the work.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the London-SIT-class rough-sleeping data-consolidation layer network: 8 components and 15 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This example follows the London Strategic Insights Tool case file. It does not rebuild the actual service. It shows how one shared layer of linked records is built and read. Three separately run record systems send records to it: CHAIN street-outreach contacts, In-Form charity casework, and H-CLIC borough statutory applications. A matcher links them into one journey per person, which all 33 London local authorities read. The tool makes no decision about any individual. Users see aggregate views only, plus their own organisation's records. The delivery team states it is not a substitute for published data and reports. Any reading that implies individual risk scoring or case decisions misreads it.
- baseline
Every borough reads the same linked layer, so one error in it shows up across the whole city at once. The matcher joins records only above an 85 percent probability of a match. That threshold keeps false positives, two different people's records matched by mistake, as few as possible. The source is the project's own second-phase data protection impact assessment, version 2.0, dated 23 October 2023. A data protection impact assessment is the review of privacy risks that UK data protection law requires for high-risk processing. It was published on the London Office of Technology and Innovation site in April 2025. It reports 91 percent recall, meaning 91 of every 100 true matches are found. It concedes that 9 in 100 are missed, so "numbers subsequently appear lower in places where they should be higher." It adds that recall varies as new data of varying quality comes in. The delivery team reports these figures itself, and no false-positive rate is published. Only a Faculty blog on the site of the trade body techUK names Splink as the matching library.
- baseline
The network draws two checks the sources do not describe. The first compares the linked layer against its source records. No source describes a routine one. The miss rate is reported by the team that built the matcher, and it moves with data quality. No source describes anyone routinely reading the linked copy back against its sources. The second is an independent evaluation of the tool's effect on decisions. No source describes one. The first-year claims of the Chief Digital Officer for London give no published figures and no comparison. They say the tool helped commissioners challenge assumptions about local rough-sleeping patterns. They say it revealed previously hidden connections between street homelessness and Housing Options services. Around this gap, the process is well governed. It has a published data protection impact assessment, hiding counts of five or fewer, a minimum data set, and retention rules. The one part not independently measured is the accuracy of the linkage everyone reads.
- assumed
The privacy risks are real, and the network marks them on the pathways that enter and read identifiable records. The linked identifiers are fuzzy-matched names, National Insurance numbers, dates of birth, and phone numbers. The people whose records are linked are not notified individually. The legal grounds the partners cite are not consent. They are a public task or a legitimate interest. For sensitive data, they are substantial public interest and research. People sleeping rough do not use the tool themselves. In user testing, a user stacking filters got one view down to 1 or 2 people. The team then followed Office for National Statistics practice: any output of 5 or fewer shows as "equal to or less than 5". A first-phase review found several sensitive fields went unused. They include substance misuse, mental health concerns, pregnancy status, prison history, care-leaver history, and entitlement to welfare benefits. The project kept them, hidden from users, for planned features. What it deleted was data outside the request and unmatched service-provider records.
- baseline
Governance followed the data. Faculty built the tool and was its sole data processor until 2 February 2024. Homeless Link then took over hosting, management, and maintenance under a Greater London Authority contract. Homeless Link already runs CHAIN, one of the three sources. So one operator now holds a source system and the linked layer. Records are kept five years per person, matching how the government's rough-sleeping indicators are defined. Someone not seen sleeping rough for five years stops counting as an existing rough sleeper. If they return after that, they are treated as new. So repeat rough sleeping after a gap of more than five years is not counted as a repeat.
- assumed
People sleeping rough, and the services they do or do not receive, are not part of how this example behaves. It follows how mistakes move between institutions only. It computes no casework or commissioning outcome for any person. The harms it reads are shared across the population. They are a steady undercount across the three systems, a risk of identifying someone, and readers over-trusting the figures despite the conceded undercount. The adoption figures, 151 onboarded users and 40 to 45 organisations, come from the vendor or the project itself, not an independent audit. So does the first-year account of impact. Forecasting demand is a stated ambition, not a working feature. A count, a journey, or a trend here is a signal about institutions, never a person.
What this example does not show
Show all 4 limitations
- People sleeping rough, and the services they do or do not receive, are not shown in this example. It shows how mistakes move between the matcher, the record stores, and the organisations that read them. The tool makes no decision about any individual. Users see population trends, plus their own organisation's records. So the harm here is shared across the population: an undercount across the three systems, a risk of identifying someone, and over-trust in the figures. It is never a decision about one person, and no outcome for any person is computed.
- The 85 percent match threshold, the 91 percent recall, and the roughly 9 in 100 missed matches come from the project's own sources. These are its second-phase data protection impact assessment, version 2.0, dated 23 October 2023, and a blog co-written by LOTI and Faculty. LOTI is the London Office of Technology and Innovation. A data protection impact assessment is the review of privacy risks that UK data protection law requires for high-risk processing. The assessment was published on the LOTI site in April 2025. No independent audit checked these figures. No false-positive rate is published, and no independent evaluation of the tool's effect on decisions was located. Adoption figures, 151 onboarded users and 40 to 45 organisations, and first-year claims of impact come from the vendor or the project. They are not independently verified. This example does not promise that any real deployment is safe.
- Counts differ across the sources because they measure different things on different dates. A Faculty blog on the site of the trade body techUK counts 33 boroughs. The Chief Digital Officer's review counts 32 boroughs and City Hall, and the assessment's list of controllers counts 32 boroughs plus the City of London. Contributing organisations are counted as about 12, then 14, then 40 to 45. Some counts are service providers only, others every contributing organisation. This example says 33 London local authorities. Forecasting demand, and adding probation, health and care, or eviction data, are stated ambitions in every source, not working features.
- This is a record-linkage system, not a risk-scoring, generative, or decision-making one. It is deliberately limited to insight about groups and whole populations. Any reading that implies scoring individuals or deciding cases misreads it. This example is about the shared linked layer and the undercount every reader inherits at once. The mayoral decision that expanded CHAIN, MD3331, does not itself name the tool.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
London's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately governed systems - CHAIN street-outreach contacts, In-Form charity casework, and H-CLIC borough statutory applications - into a single rough-sleeping journey per person that is read, in aggregate form only, across all 33 London local authorities; the tool makes no individual-level determinations, and after the build vendor's data-processor contract ended on 2 February 2024 the Greater London Authority contracted Homeless Link, which also operates the CHAIN source system, for its ongoing hosting, management, and maintenance.
empirical- Trade press techUK, Rough sleeping insights tool: Using machine learning to support decision-making across London (2024) https://www.techuk.org/resource/rough-sleeping-insights-tool-using-machine-learning-to-support-decision-making-across-london.html
- Government LOTI, GLA and London Councils, Phase 2 Rough Sleeping Strategic Insights Tool DPIA (public version, v2.0, 23 October 2023) https://loti.london/wp-content/uploads/2025/04/Phase-2-Rough-Sleeping-Strategic-Insights-Tool-DPIA-public.pdf
- Government London Office of Technology and Innovation (LOTI), Rough Sleeping Insights Project (2023-2025) https://loti.london/projects/rough-sleeping-insights-project/
The Strategic Insights Tool's matcher accepts an association only above an 85% probability threshold chosen to minimise false positives, and the project's own Phase 2 Data Protection Impact Assessment reports 91% recall - conceding that roughly 9 in 100 true cross-system matches are missed so that 'numbers subsequently appear lower in places where they should be higher' and that recall varies as new data of varying quality is ingested; no false-positive rate is published, the accuracy figures are self-reported by the delivery team, and no independent evaluation of the tool's decision impact exists.
empirical- Government LOTI, GLA and London Councils, Phase 2 Rough Sleeping Strategic Insights Tool DPIA (public version, v2.0, 23 October 2023) https://loti.london/wp-content/uploads/2025/04/Phase-2-Rough-Sleeping-Strategic-Insights-Tool-DPIA-public.pdf
- Government LOTI (Anna Humpleby) and Faculty (James MacTavish), Using AI to better understand and support homelessness interventions in London (2025) https://loti.london/blog/ai-to-better-understandhomelessness-interventions-in-london/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Housing & homelessness services domain page.
Levers available here and the patterns behind them
- Check copied records — Reconcile copied records
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Peer sharing rules — Peer-edge governance
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Store less data — Data minimization
- Gate record entries — Human-in-the-loop write gating
- Keep prompts neutral — Framing and mirroring reduction
- Upgrade model — Improve the model
Documented case histories
- London's Strategic Insights Tool: one linked memory of rough sleeping read by every borough
- Allegheny Housing Assessment
- VI-SPDAT
- LA's coordinated-entry triage revision: the fix that needed fixing
- LA County Homelessness Prevention Unit
- Santa Clara County Homelessness Prevention System
- Homebase Risk Assessment Questionnaire
- Xantura OneView (predictive homelessness flagging)
- CHAI (chronic-homelessness prediction)
- Calgary Drop-In Centre: interpretable screening a shelter's own staff choose to check
- San Jose's camera car: a low-precision detector aimed at who is sleeping outside
- Imagine LA Benefit Navigator copilot
- SafeRent Tenant Screening Score
- CrimSAFE criminal-record tenant screening
- One engine, many rivals: a shared rent-setting model and the record it writes back