Skip to content

PAN Lab example

Caddy adviser copilot at Citizens Advice

The gate is a job rather than a habit: a supervisor-checked adviser copilot

Caddy drafts benefits answers for Citizens Advice advisers, and a supervisor checks every draft before an adviser sees it. Can that check survive wider use?

See more

Caddy is an AI copilot that drafts answers to benefits questions for advisers at Citizens Advice in England and Wales. Citizens Advice Stockport, Oldham, Rochdale & Trafford built it, then scaled it with the UK Government's Incubator for AI. Caddy sends every draft to a supervisor before the adviser sees it.

How it is used

A client brings a benefits question to an adviser. Many advisers are trainees and volunteers. The adviser asks Caddy, which searches two trusted sources: GOV.UK, the government's website, and AdviserNet, Citizens Advice's internal knowledge base. Caddy drafts a short answer, about 400 words, with citations.

The adviser does not see the draft. An experienced supervisor approves, edits, or rejects it first. The adviser then passes the answer to the client in their own words.

The client only ever talks to the human adviser. Caddy makes no decision about whether anyone is eligible for, or entitled to, a benefit.

Why it was built this way

The team decided any AI tool "cannot be client-facing." They wanted to keep vulnerable, upset, or confused clients out of an automated "chatbot loop of hell."

Instead they placed the tool where advisers ask supervisors for help. The team saw that step as the organisation's biggest pain point, after remote work made timely answers from supervisors hard to get.

Citizens Advice Stockport, Oldham, Rochdale & Trafford says the supervisor's check keeps the tool within advice-accreditation standards, because Caddy itself never advises.

How it was built

The Stockport, Oldham, Rochdale & Trafford team tried and dropped an earlier chatbot in 2019. It built the first version of Caddy in about four weeks in 2023. It then improved it over about twelve months of volunteer-driven work.

The Incubator for AI sat in the Cabinet Office, and later in the Department for Science, Innovation and Technology. It announced the tool on 28 March 2024, starting at two Greater Manchester centres. The code was released as open source under the MIT licence.

The original code repository, last updated in November 2024, was archived on 21 October 2025. No public repository for Caddy 2.0 has been found.

What the trial found

The Incubator for AI describes its evaluation as a randomised controlled trial. Over about three months, more than 1,000 adviser requests for help were invisibly assigned at random. Each went either to Caddy, whose draft a supervisor then checked, or to the old route, where a supervisor answered directly. Advisers in both groups answered a survey in the chat after every call.

In a March 2025 blog post, the developers reported that about 80% of Caddy's drafts were good enough to pass to advisers without changes. Answers came back in about four minutes, roughly half the earlier ten or so. Up to 60% of supervisor time was saved per query.

Advisers with Caddy were more than twice as likely to report confidence giving advice. They were more than 1.5 times as likely to report resolving the client's issue. The government's AI Knowledge Hub states these as "twice" and "1.5 times," without "more than."

New and trainee advisers felt more comfortable asking Caddy extra questions than asking a busy supervisor. That changed how advisers asked for help, not only how fast they got it.

How far the figures go

Every figure comes from the builders, the Incubator for AI and the Stockport, Oldham, Rochdale & Trafford team, in blog posts and conference talks. No peer-reviewed trial report, preregistration, or raw data was found. The Incubator itself wrote that the trial "wasn't without issues."

An independent write-up by Stanford's Legal Design Lab describes the six-office evaluation as a four-to-six-week pilot. It does not describe how requests were assigned at random.

The confidence and resolution figures are advisers' own reports. They do not measure client outcomes or whether answers were right. No error rate after the supervisor's check has been published.

The drafts supervisors edited or rejected

The team sorted the other 20%, the drafts supervisors edited or rejected, into two kinds. Either the adviser's question lacked information, or the corpus of two sources lacked coverage.

Each kind has its own repair. For the first, it is an AI assistant to help advisers write fuller questions. For the second, it is adding material to the corpus by agreement with other organisations, for example the Child Poverty Action Group. The team expects these repairs to raise the approval rate.

Review around the pilot

Manchester Metropolitan University's People's Panel for AI, run with Manchester City Council, reviewed the tool as a panel of citizens. The team reported that the panel gave "a ringing endorsement." The pilot's governance also included what the sources call consequence scanning.

Where it is going

By 2026 the Incubator for AI listed Caddy as live. Its page says Caddy supports more than 40 local offices, with a 90% approval rating from advisers. That is advisers' rating of the tool, not a share of drafts. More than 70 offices were on a waiting list, and more than 100 signed up for the open beta of Caddy 2.0, described below. These office and approval figures are unaudited developer claims, and each count belongs to its own date.

The stated aim is full rollout to all of the roughly 250 Citizens Advice offices in England and Wales. That number is the Incubator's own framing. Citizens Advice reported helping more than 2.5 million people one to one in 2022-23. A Stanford account gives different figures: 270 local organisations at 2,540 locations, advising 2.8 million people in 2024.

Caddy 2.0 was announced on 4 November 2025, with an open beta the following month. It adds a verification engine that breaks each answer into claims and checks them against the trusted sources, "to streamline the supervisor checks." It learns from drafts supervisors reject. It also strips personal information from the adviser's question. National rollout was aimed for 2026.

In March 2025 the Incubator said it intended to try Caddy in frontline pilots in UK government departments. Third-party write-ups name immigration and tax. No 2026 source read for this case confirms those pilots are under way.

What this case asks

Unlike most cases here, this one starts with a careful design: checking is a separate person's whole job, not a habit someone under pressure might skip.

The Lab's two US benefits copilots, Nava's chatbot and the Imagine LA Benefit Navigator, leave the check to the person who uses the answer. Caddy sends every draft to a supervisor first. The case file calls that a real gain.

But it moves the risk rather than removing it. The question becomes whether Citizens Advice keeps requiring that every draft pass a supervisor. The sources show the signs. Caddy 2.0 aims "to streamline the supervisor checks." The Incubator for AI's 2026 page on Caddy describes routing answers "for human checks when needed."

Each step can be defended. Together they could turn a check on every draft into a check on some. The case file notes that a written policy changes in the open, so someone reading it closely can govern it.

What this network is drawn from

This network follows the pattern the case file describes. It is not a reconstruction of the actual tool. It shows Caddy, its search of the two sources, the corpus, the reviewing supervisors, and the frontline advisers. Clients are outside the network.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.

This case has a budget of 7 units. Each tool costs the same at every target level.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets are met before any tool is used or pressure added. Every combination of tools within the budget also keeps them met.

Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, among other targets. Before any tool is used, two failure pathways are open: Drafts sent to the supervisor, and Approved answers to advisers.

Each has one tool on offer that closes it. Escalate checks, which raises checking when monitoring flags trouble, closes Drafts sent to the supervisor. Peer sharing rules, which sets rules for what people pass to each other and makes peer review routine, closes Approved answers to advisers.

Together the two meet the targets for 4 units, and every combination within the budget that meets them includes both. Lingering effects is a Dynamics setting in which damage outlasts its cause. It is on by default, and always on under All Governance Targets. With it on, Peer sharing rules works at reduced strength. That is because Understand the system, the tool that pays for ongoing study of what the deployment is doing, is not offered here. It still closes its pathway.

More is not better here. Using every tool on offer, each at its strongest setting, costs 36 units, more than five times the budget. It meets the targets at none of the three levels that set them, because Caddy then adds too little to the work.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Caddy-class supervisor-gated adviser copilot network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 6 assumptions
  • assumed

    This network follows the pattern the Caddy case file describes: a copilot whose every draft a separate supervisor checks before the adviser sees it. It does not reconstruct the actual tool. It is a deliberate sibling of the Lab's two US benefits copilots, Nava's chatbot and the Imagine LA Benefit Navigator. In those, the person who uses the answer is meant to check it.

  • baseline

    The network treats the supervisor's review as the main check that holds mistakes back, because it is a dedicated person's job rather than a personal habit. The developers report that about 80% of drafts were approved without changes, so about one in five was edited or rejected. That figure is a result of their trial, not a measure of how often each draft goes wrong.

  • assumed

    The corpus is a closed set of two sources, GOV.UK and AdviserNet. Caddy reads it and does not write to it, so by design few mistakes enter it. The network assumes a mistake gets into the work only when a supervisor approves a subtly wrong draft or an adviser relays one. Besides the scraper's copies of the two sources, the corpus grows only through slow, negotiated additions.

  • baseline

    The network draws advisers passing approved answers to each other, and confidence in Caddy spreading across the office. The developers' trial found advisers with Caddy more than twice as likely to report confidence, and trainees more comfortable asking Caddy than a busy supervisor. The network also draws two checks that hold mistakes back: advisers flagging answers to each other, and each draft's citations.

  • assumed

    The network keeps a possible link from Caddy straight to advisers, skipping the supervisor, so you can see where Caddy would send drafts if the check were relaxed. The sources describe no such route. Caddy 2.0, announced in November 2025, aims "to streamline the supervisor checks." The Incubator for AI's 2026 page on Caddy describes routing answers "for human checks when needed." Wider use across Citizens Advice adds to that pressure.

  • assumed

    Clients are outside the network. The harm this case watches is a wrong answer that a supervisor approves and an adviser passes to a client. The Lab does not measure it. Caddy makes no decision about anyone's benefits. Every headline figure comes from the developers and has not been independently repeated. The network computes no difference in harm between groups of clients.

What this example does not show

Show all 3 limitations
  • This example does not show clients or the advice they end up receiving. The network shows how mistakes pass between Caddy, the staff, and the corpus. The sources read for this case report no measure of client outcomes.
  • Every headline figure comes from the tool's builders, in blog posts and conference talks, and none has been independently repeated. That covers the 1,000-plus requests assigned at random, the roughly 80% of drafts approved without changes, the four-minute turnaround, and the doubled adviser confidence. The confidence and resolution figures are advisers' own reports, not measures of client outcomes or accuracy. No error rate after the supervisor's check has been published. One independent account describes the evaluation as a short pilot, without detailing how requests were assigned at random.
  • A safe start in this network is not a safety promise for any real deployment. On the sources' account, the check on every draft that makes this network safe is already under pressure to relax as use spreads.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In a developer-reported randomised controlled trial of more than 1,000 adviser support requests, an adviser-facing benefits copilot at Citizens Advice returned supervisor-checked answers in about four minutes, roughly half the previous response time, with about 80 percent of its drafts approved by supervisors without revision; these figures are reported by the tool's builders and have not been independently replicated.

    empirical
    • Government Varotsis, Transforming Civic Engagement with Caddy (Incubator for Artificial Intelligence, i.AI, UK Government, developer blog, 2025) https://ai.gov.uk/blogs/transforming-civic-engagement-with-caddy/
    • Government Department for Science, Innovation and Technology / i.AI / CASORT, Caddy (AI Knowledge Hub use case, 2025) https://ai.gov.uk/knowledge-hub/use-cases/caddy/
    • Academic Stanford Legal Design Lab / Justice Innovation, How AI is augmenting human-led legal advice at Citizens Advice (Caddy adviser copilot) https://justiceinnovation.law.stanford.edu/how-ai-is-augmenting-human-led-legal-advice-at-citizens-advice/
  • Advisers given access to the copilot were reported to be more than twice as likely to say they felt confident giving advice than a control group, a self-reported measure from post-call in-chat surveys rather than a client-outcome or accuracy measure.

    empirical
    • Government Varotsis, Transforming Civic Engagement with Caddy (Incubator for Artificial Intelligence, i.AI, UK Government, developer blog, 2025) https://ai.gov.uk/blogs/transforming-civic-engagement-with-caddy/
    • Academic Stanford Legal Design Lab, Caddy Q and A copilot (JusticeBench project page, 2025) https://www.justicebench.org/project/caddy

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories