PAN Lab example
Nava assistive benefits chatbot
Done carefully: a verify-before-use copilot
Nava's chatbot answers caseworkers' benefits questions from vetted documents, with quotes to check. The question is what keeps this cautious design in place under pressure.
See more
Nava Labs, part of Nava Public Benefit Corporation, built and piloted an AI chatbot for benefits caseworkers. Caseworkers ask it questions during calls with clients, and it summarizes an answer from a set of vetted policy documents. It attaches direct quotes it does not alter, so the caseworker can check the answer before passing it on.
How it is used
A client on the phone asks a caseworker about public benefits. The caseworker asks the chatbot, reads its answer, and checks the quoted source. The design has the caseworker do that check before passing the answer to the client.
Nava positions the chatbot as help for benefit navigators, the caseworkers who help people apply, not as an authority for applicants. A professional stands between the chatbot and the person applying.
Why the design is cautious
The case file places this chatbot at the cautious end of the ways such a chatbot can be designed. Its scope is bounded, and it answers from vetted policy sources. Nava published its reasoning for the design, which makes outside scrutiny possible.
Nava describes the document set as predefined and vetted. The sources read for this case do not describe the chatbot writing into it.
The case file contrasts it with public-facing chatbots that answered with confidence and got it wrong. One told a business to break the law. The underlying technology is the same. The case file says what differs is the design around it.
What this case asks
Unlike most cases here, this one starts with a careful design. So the usual question flips. The question is what keeps the design careful when pressure rises.
Where the facts come from
Every source for this case is Nava's own published material. They are Nava's pilot announcement with Imagine LA, its evaluation write-up, and its piece with Benefits Data Trust on AI for benefit navigators. None is an independent evaluation.
The sources name a 2025 pilot with Imagine LA, funded by a Gates Foundation grant. A later Google.org grant funds Nava's work on AI agents. The sources also list Nava work in California, Texas, and Pennsylvania, without saying what was done there.
The Lab draws the Imagine LA pilot as a separate case. The notes behind that case say this one draws Nava's earlier, exploratory work. The two sets of notes disagree.
What this network is drawn from
This network follows the pattern the sources describe. It is not a reconstruction of Nava's own system. It shows the chatbot, its search of the vetted documents, the document set, and the caseworkers. Clients are outside the network.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 7 units. Each tool costs the same at every target level.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets are met before any tool is used or pressure added. Any single tool on offer also keeps them met.
Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, among other targets. Before any tool is used, one failure pathway is open: Answers with quotes to caseworkers.
Escalate checks is the only tool on offer that closes it. When monitoring flags trouble, it has caseworkers check the chatbot's answers more closely, instead of waiting for the next review. It meets the targets on its own for 2 units. Every combination within the budget that meets them includes it.
More is not better here. Using every tool that can be combined, each at its strongest setting, costs 33 units, nearly five times the budget. It meets the targets at none of the three levels that set them, because the chatbot then adds too little to the work.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Nava-class verify-before-use copilot network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 4 assumptions
- assumed
This network follows the pattern the Nava case file describes: a chatbot that answers from vetted documents, with quotes a person checks before use. It does not reconstruct Nava's actual product.
- assumed
The network draws caseworkers passing answers to each other. It also draws two checks that hold mistakes back: caseworkers check each other's answers, and the direct quotes let each answer be checked against its source. It keeps a possible link from the chatbot to another AI system on the map, so you can see a pressure open it. The sources do not describe caseworkers sharing answers or checking each other's answers.
- baseline
The network assumes the chatbot writes little into the document set, so few mistakes enter the set before any tool is used or pressure added. Nava describes the set as predefined and vetted. The sources read for this case do not describe the chatbot writing into it.
- assumed
The network assumes a mistake gets into the work only when a caseworker uses an answer without checking its quoted source. That is the pathway this case watches: Answers with quotes to caseworkers.
What this example does not show
Show all 1 limitation
- This example does not promise that any real deployment is safe. Its network starts with few mistakes passed between its parts. That start comes from Nava's own descriptions of the chatbot, and none of its sources is an independent evaluation.
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Keep prompts neutral — Framing and mirroring reduction
- Keep skills sharp — Deskilling-arrest mandate
- Gate vendor updates — Vendor quality gate
- Store less data — Data minimization
- Peer sharing rules — Peer-edge governance
- Assign a challenger — Structured dissent
- Escalate checks — State-feedback vigilance
Documented case histories
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down