PAN Lab example
Imagine LA Benefit Navigator copilot
Best where you can check it least: a benefits-navigation copilot
A chatbot answers Los Angeles caseworkers' benefits questions, quoting policy. Accuracy rose most for new staff and hard questions, where errors are hardest to spot.
See more
Nava Labs, part of Nava Public Benefit Corporation, built and piloted a generative AI chatbot on Imagine LA's Benefit Navigator platform. Caseworkers type a client's question during a call. The chatbot answers from approved government documents and attaches direct quotes, so the caseworker can check the answer before relaying it.
How it is used
Benefit Navigator is Imagine LA's digital tool for benefit navigators, staff who help families apply, and for case managers. They use it to connect low-income Los Angeles County families to public benefits. Imagine LA is a 501(c)(3), a tax-exempt nonprofit, with more than eighteen years of direct services. Amplifi handled the platform integration, the technical work of fitting the chatbot into the platform.
The chatbot answers questions about eligibility, application steps, documentation, and how programs interact. It covers roughly forty Los Angeles County benefit programs and tax credits. It answers in several languages, and it can restate an answer in simpler language or in Spanish.
It is assistive. Caseworkers use it while on the phone with a client. They check its answer against the quoted source before relaying it. It does not make eligibility decisions on its own.
Who paid for it
A $3.6 million Gates Foundation grant, announced in December 2024, funded the pilot. Nava's announcement says the grant also covers tools for document processing, referral generation, and case-note summaries.
A $1.5 million Google.org grant in June 2025 funds Nava's next step, with Imagine LA and First 5 Riverside County. It is a move toward AI agents. They would navigate databases and benefit portals, and complete applications under caseworker supervision.
How it was evaluated
Nava and its research partners, Cornell University and Georgetown University's Better Government Lab, ran the evaluation. It had two parts.
The first was a randomized controlled trial, a study that assigns people to groups at random so the groups can be compared. It had 125 caseworkers. It measured how accurate their answers to hypothetical client questions, built from real experiences, were when they were shown the chatbot's answers.
The second was a fourteen-week pilot in real work. It had 61 caseworkers across six organizations in Los Angeles County.
What the evaluation found
Using the chatbot was estimated to improve answer accuracy by about 40 percent on average. The gains were largest on the hardest questions and among newer, less-experienced staff. The evaluation gives that as a direction, not a measured breakdown.
About 65 percent of caseworkers with access used the chatbot, submitting about 14 prompts, or typed requests, each on average. Satisfaction was modest and low-positive: a Net Promoter Score of 11, with 40 percent classified as promoters. That score is a common measure of how likely users are to recommend a product. Use tended to decline over time and varied by site. The evaluation reported that sustained use needs ongoing support from the organizations.
Answers averaged a tenth-to-twelfth-grade reading level. That is easier than the college-level source manuals, but short of a plain-language goal.
Results on administrative burden and time savings were promising but inconclusive. They were not statistically significant, because of sample-size limits and a lower response rate.
Who wrote the evaluation
Nava shared the pilot results at a Demo Day in December 2025. The work was published in March 2026 as a Nava case study. The same team published it as the academic report "Helping the Helpers".
Its authors are Allison Koenecke and Jennah Gosciak of Cornell, Eric Giannella and Zhaowen Guo of Georgetown's Better Government Lab, and Michael Chen and Martelle Esposito of Nava. So the tool's builder co-wrote the headline result with its academic partners. No fully independent replication exists.
Figures that belong to the platform
Some figures often cited nearby belong to the Benefit Navigator platform, not to the chatbot. An earlier platform pilot, from September 2023 to November 2024, reached more than 500 case managers. It assisted over 10,000 beneficiaries and secured on average an additional $10,869 in benefits per household.
What the case file draws from it
The case file argues that the measured benefit and the risk sit in the same place. The chatbot helped most on the hardest questions and for the newest staff. Those are the questions where a wrong answer is hardest to spot, and the staff least able to catch it.
The whole safeguard is one habit: read the quoted source before relaying the answer. The evaluation shows that habit is not settled. Use reached about two-thirds of caseworkers, then declined without sustained engagement. The case file adds that the fluent answers can sound more certain than the policy behind them.
The chatbot flags no one and denies no one. So the case file sees the governance work as keeping a habit alive. It warns that the funded move toward agents that complete applications under supervision would take away the check a caseworker makes on the phone.
How this case differs from the Nava case
The Lab also has a case on Nava's chatbot for benefits caseworkers, drawn from Nava's earlier, exploratory work. This case draws the Imagine LA pilot, which put the same careful design to a randomized test. The test showed where the gains landed and that use faded. The earlier case could not show either.
What this network is drawn from
This network follows the pattern the case file describes. It does not reconstruct the actual chatbot. It shows the chatbot, its search of the approved documents, the documents, the caseworkers, and the evaluation. The families served are outside the network.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 7 units. Each tool costs the same at every target level, except Understand the system, which pays for ongoing study of the deployment. It costs 3 units under Explore (No Targets) and Service Targets Only, and 4 under the two higher levels. While it is on, Escalate checks, Know the tool, Keep prompts neutral, and Vet connections each cost 1 unit less. With its stronger setting, Deep research, which costs 6 units, they cost 2 less, but never less than 1.
Before any tool is used, one failure pathway is open: Quoted answers to caseworkers. The network is self-correcting, meaning it clears mistakes on its own. The caseworkers are keeping up with the work.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets are met before any tool is used. Every set of tools within the budget keeps them met.
Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, among other targets. Escalate checks is the only tool on offer that closes Quoted answers to caseworkers. It meets the targets on its own for 2 units. Every set of tools that meets them includes it. Within the budget, 45 different sets of tools meet Service and Safety Targets, and 44 meet All Governance Targets.
The two higher levels also ask you to refill the Privacy gauge, the Lab's reading of how well client data is kept safe. It starts below full. With no pressure added, no pathway here drains it, so it refills as the clock runs.
More is not better. Using every tool, each at its strongest setting, costs 29 units, more than four times the budget. It meets Service Targets Only and Service and Safety Targets, but not All Governance Targets. At that level, the chatbot then adds too little to the work.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Benefit-Navigator-class assistive copilot network: 5 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
This network follows the pattern the Imagine LA Benefit Navigator case file describes: a chatbot that answers from approved documents, with quotes a caseworker checks before use. It does not reconstruct the actual chatbot. The Lab also has a Nava case, drawn from Nava's earlier, exploratory benefits work. This case draws the later Imagine LA pilot, funded by the Gates Foundation and evaluated with academic partners. Its gains were largest for the least-experienced staff.
- baseline
The evaluation found the largest accuracy gains on the hardest client questions and among the newest, least-experienced staff. So the network assumes caseworkers lean on the chatbot's answers from the start. The case file notes that those who lean hardest are the least able to catch its mistakes.
- baseline
The network assumes safety here is a habit that must be kept up, not a fixed feature of the design. About 65 percent of caseworkers with access used the chatbot, and use declined over time without sustained engagement. So the network treats the checking habit, and the engagement behind it, as able to fade.
- assumed
The approved documents are a curated set of government documents kept by people. The published sources on the pilot describe no writing by the chatbot into them. So the network assumes few mistakes enter the documents before any tool is used or pressure added. A mistake enters the work mainly when a caseworker uses an answer without checking its quoted source.
- assumed
The network draws caseworkers passing answers and practice to each other across the six pilot sites. The sources do not describe this. It also draws the direct quotes as a check a person can make on each answer. It keeps a possible link from the chatbot to another AI agent on the map. The funded move toward agents that navigate portals and complete applications under supervision could open it.
- assumed
The network draws the evaluation as a periodic check on the chatbot, not a second read of each answer. The tool's builder co-wrote it with its academic partners, and no fully independent replication exists. The reported gain of about 40 percent in accuracy comes from a randomized trial on hypothetical client questions. It is not an audit of live eligibility decisions.
- assumed
The families served are not part of this network. The harm here would be a wrong or missing answer passed to a client, recorded outside a network like this one. The chatbot flags no one and denies no one. No difference in harm between groups of clients is computed here. The answers' reading level, tenth to twelfth grade against college-level manuals, is an accessibility caveat, not a gap between groups.
What this example does not show
Show all 2 limitations
- This example does not show the families served, or which benefits they do or do not receive. The Lab shows only how mistakes can move among the chatbot, its documents, the caseworkers, and the evaluation. Any outcomes for families are recorded in the case file and measured outside a network like this one.
- The accuracy gain of about 40 percent comes from a randomized trial on hypothetical client questions. The tool's builder co-wrote it with its academic partners, and no one has replicated it independently. It measures the accuracy of caseworkers' answers, not whether more families enrolled. Reported time savings were promising but inconclusive.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The same evaluation reported that the chatbot's accuracy gains were largest on the most difficult client questions and among the newest, least-experienced staff (a directional finding, not a quantified breakdown), that about 65% of caseworkers with access used it at an average of about 14 prompts each and a modest, low-positive satisfaction (a Net Promoter Score of 11), that usage tended to decline over time without sustained engagement, and that answers averaged a tenth-to-twelfth-grade reading level against college-level source manuals.
empirical- Vendor Nava Public Benefit Corporation (Nava Labs), Evaluating a GenAI-powered assistive chatbot for caseworkers (2026) https://www.navapbc.com/case-studies/evaluating-ai-assistive-chatbot-caseworkers
- Academic Chen, Esposito, Giannella, Guo, Gosciak, Koenecke, Helping the Helpers: Evaluating a GenAI-powered assistive chatbot for caseworkers (Georgetown University Better Government Lab, Cornell University, and Nava PBC, 2026) https://digitalgovernmenthub.org/library/helping-the-helpers-evaluating-a-genai-powered-assistive-chatbot-for-caseworkers/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Housing & homelessness services domain page.
Levers available here and the patterns behind them
- Keep skills sharp — Deskilling-arrest mandate
- Escalate checks — State-feedback vigilance
- Know the tool — AI literacy & boundary rules
- Keep prompts neutral — Framing and mirroring reduction
- Mark AI-written records — Provenance labeling
- Gate vendor updates — Vendor quality gate
- Review on schedule — Oversight cadence & retrospectives
- Vet connections — Connection authorization
- Understand the system — Understand the system
- Upgrade model — Improve the model
Documented case histories
- Imagine LA Benefit Navigator copilot
- Allegheny Housing Assessment
- VI-SPDAT
- LA's coordinated-entry triage revision: the fix that needed fixing
- LA County Homelessness Prevention Unit
- Santa Clara County Homelessness Prevention System
- Homebase Risk Assessment Questionnaire
- Xantura OneView (predictive homelessness flagging)
- London's Strategic Insights Tool: one linked memory of rough sleeping read by every borough
- CHAI (chronic-homelessness prediction)
- Calgary Drop-In Centre: interpretable screening a shelter's own staff choose to check
- San Jose's camera car: a low-precision detector aimed at who is sleeping outside
- SafeRent Tenant Screening Score
- CrimSAFE criminal-record tenant screening
- One engine, many rivals: a shared rent-setting model and the record it writes back