PAN Lab example
GOV.UK Chat
The gate that said not yet: a public assistant behind a staged pilot gate
GOV.UK Chat answers the public's questions from official guidance. Its gate held back a version that fell short, but its builders grade its accuracy themselves.
See more
GOV.UK Chat is an AI assistant built and run by the Government Digital Service (GDS), part of the Department for Science, Innovation and Technology. People ask it about tax, benefits, visas, and driving in the GOV.UK app. A commercial language model writes each answer only from official GOV.UK guidance, with links back to the source pages.
How it is used
A person types a question in their own words. The question is matched against about 700,000 passages of GOV.UK guidance, with no open web search. The language model, hosted on a commercial cloud platform, writes an answer only from the passages found. It is told to ignore what it learned in training.
A safety check screens each answer. When a question is unclear, the assistant asks a clarifying question. Every answer links to its GOV.UK source pages, with a reminder to check them.
It gives information and makes no decision. GDS states it "does not attempt to provide advice" and "clearly signposts when users should check the original guidance." No one's eligibility, entitlement, or sanctions are decided by it.
How it was released, stage by stage
GDS's first generative AI experiment, in 2023, went through internal red teaming, where testers try to make it fail. A test with about a dozen users followed, then a private test with 1,000 invited users. In a survey of 157 of them, nearly 70 percent found the answers useful and just under 65 percent were satisfied.
In findings published January 18, 2024, GDS reported that answers "did not reach the highest level of accuracy demanded for a site like GOV.UK." There were "a few cases of hallucination," answers the AI made up. GDS held that version back and called its approach "not moving fast and breaking things."
In November 2024, GDS opened a private beta with a waiting list, on selected GOV.UK business pages. It cited a consistent improvement in accuracy since 2023. On October 7, 2025, it published a formal Algorithmic Transparency Record describing how the system works.
Two public pilots ran on GOV.UK web pages from late 2024 and in the GOV.UK app in autumn 2025. The app had entered public beta in July 2025. GOV.UK Chat became available to all app users on March 26, 2026, and officially launched on May 14, 2026. As of mid-2026, GDS was still exploring whether to offer it on the GOV.UK website.
What GDS reports
The web pilot drew 10,136 users asking 23,838 questions. The app pilot drew 641 users asking 2,670 questions in four weeks. GDS called the two pilots "the government's biggest public test of generative AI to date."
GDS reports measured accuracy rising from 76 percent at its earliest benchmark, the first accuracy measurement it reports, to 90 percent by the app pilot. Its evaluation team's subject-matter experts and automated measures scored it. After GDS added clarifying questions, the assistant answered 88 percent of questions within its scope. GDS reports 508 attempts to jailbreak it during the pilots, tricks to make it break its rules. Its safety guardrails stopped all of them.
In a follow-up survey of app users, 73 percent found it useful and 64 percent were satisfied. Answers took 10.7 seconds on average, and satisfaction rose when faster answers were simulated. GDS noted the "latest versions of frontier models have been more powerful but slower," and chose accuracy over speed.
By the launch, more than 7,800 people had asked more than 15,000 questions since the soft launch. Demand was strongest for tax, driving and transport, and benefits.
Who grades the numbers
Nearly every figure here is reported by GDS itself. The sources describe no independent audit of how accuracy is scored. How many answers were scored, and how they were chosen, is unpublished.
The one outside evaluation in the sources is a security test for jailbreaks, run with the AI Security Institute, a UK government body. It is not an accuracy audit. GDS also claims that, on government questions, GOV.UK Chat scores higher than widely used consumer AI assistants. That is GDS's own comparison. Civil Service World, a trade publication, reported the same accuracy record in March 2026.
The guidance behind the answers
In December 2025, GDS wrote that GOV.UK Chat "can only be as good as the content published on GOV.UK by departmental teams. Clear and accurate content will result in better answers." The case file reads this as a loop: failed answers become pressure on departments to fix their guidance.
GDS describes Chat joining guidance across departments. A new parent asking "what help can I get?" needs answers from HMRC, the tax authority, the Department for Work and Pensions, and the Department for Education. The sources say the assistant does not write its answers into the guidance.
GOV.UK Chat was selected as a Prime Minister's AI Exemplar. The sources do not explain what that selection involves. GDS says its team is experimenting with agentic AI, AI that carries out tasks, for future transactions.
How this case compares
The Lab's case on New York City's MyCity business chatbot has the same shape. It was a public-facing city chatbot that people asked directly, with no professional in between. Its errors came in an official voice.
MyCity's errors were published in an investigation and again in a formal audit. The city kept it online and rejected all seven audit recommendations, until an unrelated budget cut ended it. GOV.UK Chat was built with a gate that could say "not yet", and GDS used it.
What this case asks
The question is not whether the gate exists. It does. The question is whether anyone outside GDS checks the accuracy figure the gate relies on.
The quieter risk is the one GDS itself named: misplaced confidence. A fluent answer can be taken as settled without a click back to the source.
What this network is drawn from
This network follows the pattern the sources describe. It is not a reconstruction of GOV.UK Chat itself. It shows the assistant, the staged release gate, the members of the public, the GDS evaluation team, and the GOV.UK guidance collection. It also shows a pathway out of the service's controls.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 9 units. Each tool costs the same at every target level.
The work here is answering the questions people bring. Explore (No Targets) sets no targets. Under Service Targets Only, the targets ask that mistakes stop building on one another, and that GOV.UK Chat keep helping the work. Before any tool is used, mistakes still build on one another. Any single tool on offer meets the targets.
Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, among other targets. Before any tool is used, six are open.
Three carry the assistant's answers: to the public, to the evaluation team, and to the release gate. Two carry the guidance: to the assistant, and to the public through the source links. The sixth carries failed answers back to the guidance.
Every combination within the budget that meets these targets includes Escalate checks and Mark AI-written records. Escalate checks closes the three answer pathways. Mark AI-written records closes the two guidance pathways. Gate record entries or Store less data, at 3 units each, closes the sixth. Gate record entries requires sign-off before changes enter the guidance collection. Store less data has less written into the collection and kept, so fewer mistaken changes go in. The cheapest combinations cost 7 units.
All Governance Targets asks more of the service. Gate record entries slows the work more than Store less data does. So adding Review on schedule to Escalate checks, Mark AI-written records, and Gate record entries leaves GOV.UK Chat helping too little. With Store less data in place of Gate record entries, the four meet the targets.
More is not better here. Using every tool, each at its strongest setting, costs 38 units, over four times the budget. It meets the targets at none of the three levels that set them, because GOV.UK Chat then adds too little to the work.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the GOV.UK-Chat-class staged-gate public assistant network: 6 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
This network follows the pattern the GOV.UK Chat case file describes: a public AI assistant released in stages behind an accuracy gate. It does not reconstruct the actual product or any of its releases. It is a contrast with the Lab's case on New York City's MyCity business chatbot. In both, the public asks directly, with no professional in between, and errors come in an official voice. MyCity stayed online after its errors were documented, and the city rejected its audit. Here the gate could say "not yet", and did. It also differs from the Lab's two cases on Nava Labs' benefits chatbots, one piloted with Imagine LA. Those chatbots answer caseworkers, who check before telling the person. Here the answer goes straight to the public.
- baseline
The network draws the staged release gate as a working check, not a badge. GDS scored answers for accuracy at each stage before a wider release. It held back the 2023 version because answers fell short of the accuracy demanded for a site like GOV.UK. There were a few made-up answers. The case file records which stages passed or were held. The network does not compute that. The gate is the heart of this case. The pathway it is meant to protect is the answers going to the public.
- baseline
The network keeps an independent accuracy check on the map, so you can see a tool add it. The sources describe no independent audit of how accuracy is scored before a release. Nearly every reported figure is GDS's own. That covers accuracy rising from 76 to 90 percent, the 88 percent answer rate, and the 508 blocked jailbreak attempts. How many answers were scored, and how they were chosen, is unpublished. The one independent evaluation in the sources is a security test for jailbreaks, run with the AI Security Institute, a UK government body. It is not an accuracy audit.
- baseline
The network draws the guidance collection as protective. That is unlike the Lab's cases of workplace AI assistants that save their own drafts into shared records. The sources say GOV.UK Chat does not write its answers into the official guidance. The network keeps that pathway on the map, so you can see a pressure open it. Failed answers become pressure on departments to fix their guidance. GDS says Chat "can only be as good as the content published on GOV.UK." Every answer sends the reader back to the official source page rather than replacing it.
- assumed
How often GOV.UK Chat makes a mistake in one answer is a modeling choice, not a measured rate. The reported rise from 76 to 90 percent accuracy is GDS's own figure. Experts and automated measures scored it over pilot samples. The network sets the assistant's mistakes lower than for public chatbots released without a gate. Its answers come only from curated official guidance, it is told to ignore its training, and releases pass an accuracy gate.
- assumed
The network assumes ways of scoring and testing spread among the evaluation team's members, and that members check each other's scores. It also assumes that one assistant answering everyone makes a wrong answer repeat for everyone who asks, rather than scatter. These are modeling assumptions, not measurements.
- assumed
The members of the public appear in this network only as people who use the answers. No benefit, harm, or outcome to any person is computed here. The sources hold no error rate by topic, such as benefits, and no complaint data. GDS states the system does not attempt to provide advice, and it points people to the original guidance. What an answer means for the person who acts on it is covered in the case file. It is measured outside this network. The network shows only how mistakes pass between the assistant, the people, and the guidance.
What this example does not show
Show all 3 limitations
- Nearly every figure in this case is reported by GDS, the system's operator. That covers the 76 to 90 percent accuracy, the 88 percent answer rate for questions within scope, the 508 blocked jailbreak attempts, and the satisfaction rates. The sources describe no independent audit of how accuracy is scored. How many answers were scored, and how they were chosen, is unpublished. The claim that it outscores consumer AI assistants on government questions is GDS's own comparison. Treat these figures as the operator's claims, not independent measurements.
- GOV.UK Chat gives information and makes no automated decision. It decides no one's eligibility, entitlement, or sanctions, so the appeal and override questions of cases where an automated system decides about people do not apply. The people who use it appear here only as users of the answers. There is no error rate by topic, such as benefits. No data has been published on harm, complaints, or outcomes from the answers that were not accurate. At the latest pilot, GDS measured about one in ten of the answers it scored as not accurate. What an answer means for the person who acts on it is covered in the case file, and measured outside this example.
- The documents label the phases inconsistently, as an experiment, a private beta, or a pilot. Each phase's user and question counts belong to the source that reports them. The transparency record describes a private beta capped at 2,000 users over four weeks. The pilot-completion post reports 10,136 users in the first public pilot. These describe different phases and set-ups. The claim that all 508 jailbreak attempts were blocked sits beside the transparency record's own caveat that success can never be guaranteed.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The UK Government Digital Service ran what it called the government's biggest public test of generative AI to date: across two gated public pilots (a late-2024 web pilot of 10,136 users asking 23,838 questions, and an autumn-2025 GOV.UK app pilot of 641 users asking 2,670 questions in four weeks), more than 10,000 people asked GOV.UK Chat about 26,000 questions on tax, benefits, and visas. Its first 2023 version was held back in findings published January 18, 2024 because, GDS reported, answers did not reach the highest level of accuracy demanded for a site like GOV.UK, including a few cases of hallucination. GDS reports measured answer accuracy rising from 76 percent (its earliest benchmark) to 90 percent by the autumn 2025 pilot, assessed by subject-matter experts plus automated evaluation, an 88 percent answer rate for in-scope questions after a clarifying-questions feature was added, and that 508 attempts to jailbreak the system across the pilots were all prevented by its guardrails; it soft-launched to all GOV.UK app users on March 26, 2026 and officially launched on May 14, 2026. Nearly every one of these figures is self-reported by GDS, the system's operator, and the accuracy denominators and sampling frames are unpublished.
empirical- Government evaluation Government Digital Service (Inside GOV.UK), 5 things we learned testing GOV.UK Chat: an AI assistant for government (2026) https://insidegovuk.blog.gov.uk/2026/03/16/5-things-we-learned-testing-gov-uk-chat-an-ai-assistant-for-government/
- Government evaluation Government Digital Service (Inside GOV.UK), The findings of our first generative AI experiment: GOV.UK Chat (2024) https://insidegovuk.blog.gov.uk/2024/01/18/the-findings-of-our-first-generative-ai-experiment-gov-uk-chat/
- Government Government Digital Service, Answers in seconds, 24/7: GOV.UK Chat launches in the GOV.UK app (2026) https://gds.blog.gov.uk/2026/05/14/gov-uk-chat-launches/
GOV.UK Chat is a retrieval-augmented assistant that, per its Algorithmic Transparency Record published October 7, 2025, answers only from roughly 700,000 vectorised chunks (36.9 GB) of curated official GOV.UK guidance, is instructed to ignore its training data, rejects questions containing phone numbers, emails, or card numbers, links every answer back to its GOV.UK source pages with a reminder to verify, and retains question data encrypted for 12 months; GDS states it does not attempt to provide advice and makes no automated decision. GDS's December 2025 vision post frames a content-dependency loop, stating that GOV.UK Chat can only be as good as the content published on GOV.UK by departmental teams. The record's independent evaluation is a jailbreak (security) assessment conducted with the AI Security Institute, alongside the record's own caveat that it is not possible to guarantee no jailbreaking attempts will succeed; there is no independent audit of the accuracy methodology, and GDS's claim that for government-related questions the tool scores higher than widely-used consumer AI assistants is the operator's own comparison.
empirical- Government Department for Science, Innovation and Technology (GOV.UK Algorithmic Transparency Recording Standard), GOV.UK Chat Algorithmic Transparency Record (2025) https://www.gov.uk/algorithmic-transparency-records/dsit-gov-dot-uk-chat
- Government Government Digital Service (Inside GOV.UK), GOV.UK has entered the Chat: our vision for GOV.UK Chat (2025) https://insidegovuk.blog.gov.uk/2025/12/16/gov-uk-has-entered-the-chat-our-vision-for-gov-uk-chat/
- Trade press Civil Service World (Jim Dunton), GOV.UK AI chatbot achieves 90% accuracy (2026) https://www.civilserviceworld.com/professions/article/govuk-ai-chatbot-achieves-90-accuracy
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Mark AI-written records — Provenance labeling
- Check with a second model — Cross-model verification
- Escalate checks — State-feedback vigilance
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- GOV.UK Chat
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down