Skip to content

PAN Lab example

DWP Whitemail Insights and Vulnerability Scanner

The letter no one reads twice: an upstream vulnerability scanner

The UK welfare department's AI scans about 25,000 letters daily, flagging signs of risk like self-harm for caseworkers. No one rechecks the rest for risk.

See more

The Whitemail Insights and Vulnerability Scanner is an AI tool the UK Department for Work and Pensions (DWP) runs on paper post from citizens. It reads each scanned letter and flags signs of possible vulnerability, such as suicide and self-harm or domestic violence. It sorts the letters it does not flag into routing themes, such as change of address.

How it works

The tool converts handwritten content into machine-readable text. It then uses a pre-trained, open-source language model that sorts text into categories it is given, without examples of each. The case file calls this zero-shot classification.

Each document goes to the Vulnerability Scanner first. It checks the letter against eight prescribed themes. They include suicide and self-harm, domestic violence and abuse, drugs and alcohol misuse, financial hardship, and mental health. It attaches a rationale naming the theme to each flag.

Only documents it does not flag go on to Whitemail Insights. That part sorts them into nine routing themes, such as change of address and change of bank, so they reach the right benefit lines.

Personal data is automatically redacted after scanning. The output to trained staff is an anonymised daily report of flagged customers, DWP's word for claimants, not a decision. Each scanned document carries a unique identifier. Staff can use it to find the customer's record and the original letter image, in systems separate from the two AI tools.

The system is hosted on cloud infrastructure with end-to-end encryption and no internet connectivity.

Who built it

Accenture (UK) Limited was the developer and delivery lead, working under DWP project managers. The work was part of "The Garage" partnership. PublicTechnology, a trade publication, reported that arrangement as potentially worth around £49m. DWP keeps the intellectual property.

Earlier research by Anna Dent, an independent analyst, drew on freedom-of-information requests and suspected a different supplier, Agilysis. The case file treats the attribution in the government's published transparency record for the tool as authoritative.

What DWP says it does and does not do

DWP states the tool "does not make benefit entitlement decisions or influence benefit entitlement decision-making." Staff assess each flagged situation case by case.

The Secretary of State replied to Parliament's Work and Pensions Committee on 4 December 2023. The reply committed that "any decisions that could affect the continued payment of benefits to customers are made by colleagues rather than by machines."

Richard Corbridge, DWP's Chief Digital and Information Officer, spoke to Computing, a trade publication, in March 2024. He said: "No important decision [at DWP] is made about you by any computer, it is a human that's making the decisions."

PublicTechnology described the tool as "a new and additional service" that did not replace existing manual casework processes.

How much it handles

The Secretary of State's December 2023 reply said the technology sped up finding vulnerable people among "the around 22,000 letters the department receives each day." It said the process "now takes a day rather than weeks."

In the March 2024 interview, published on 22 March, Corbridge said the tool had been live for six months. He said it was analysing 22,000 documents a day and had processed over 2 million. He said letters were sorted the day they arrived, not after weeks.

A Computing IT Leaders 100 profile of Corbridge, published on 10 June 2024, carried a further claim. It said a response used to take four to six weeks before the system went live, and that 75% now come the same day.

The transparency record, published on 27 November 2025, gives about 25,000 documents a day. These throughput and turnaround figures are claims by DWP and its ministers and officials. None has been independently verified.

How it is tested and reviewed

Testing used synthetic, real, and production-like data, with at least 15 synthetic documents for each theme. The record names three measures for judging the results: precision, recall, and F1-score. Precision measures how many of the flags are right. Recall measures how many vulnerable letters get flagged. F1-score combines the two. The record discloses no values for any of them.

After deployment, DWP reviews the outputs by hand every day. It adjusts the thresholds, the cut-off points for raising a flag, as it goes.

DWP's governance reviews included a data protection impact assessment, an Equality Analysis, a Government Internal Audit Agency review, an external assessment, and a security review. They concluded that "no key risk was identified." That is DWP's own conclusion, not an outside finding.

The 2023 correspondence with Parliament also records an AI steering board chaired by the Chief Digital and Information Officer. It records a ministerial oversight appointment, Viscount Younger of Leckie.

What critics found

Robert Booth reported in the Guardian in January 2025, using freedom-of-information requests. He found that claimants are not told the AI reads their letters. DWP's internal data protection impact assessment stated that letter writers "do not need to know about their involvement in the initiative."

The letters can include national insurance numbers, dates of birth, health information, bank details, racial and sexual characteristics, and children's details, including special needs.

The tool had been piloted since at least 2023. At the time of that reporting, it had not been logged on the central government AI transparency register, despite a ministerial mandate. The transparency record appeared roughly two years after deployment began.

Meagan Levin, policy manager at the organisation Turn2us, voiced "serious concerns." She said "transparency and accountability must be at the heart of any AI system." On the tool's core function, she said "prioritising some cases inevitably deprioritises others." She added that "it is vital to understand how these decisions are made and ensure they are fair."

Anna Dent found on 16 December 2025 that the transparency record is itself incomplete. It "refers to a section which doesn't exist." So the words or phrases DWP treats as signs of each vulnerability are still undisclosed.

Her earlier mapping, on 3 February 2025, placed this tool among DWP's other AI projects. It included two halted proofs of concept: A-cubed, which summarised policy for work coaches, and Aigent, which aimed to speed up Personal Independence Payment decisions.

The Independent reported on 27 January 2025 that at least half a dozen DWP AI prototypes had been shelved. The whitemail tool remained in use.

The error no one sees

Human oversight here is real but one-sided. A flagged letter gets a trained caseworker, who can pull the original letter image and assess the case. A letter the scanner does not flag stays in the ordinary queue. It is routed by theme and never read again for vulnerability.

Letter writers are not told the tool exists. The case file argues that a missed flag therefore brings no complaint, because no one knows there was a decision to appeal. The case file reads this as an error that surfaces, if ever, as harm blamed on something else. That reading is an inference from the documented design, not a finding.

The concern that the tool draws attention toward the flagged and away from the unflagged is voiced by Turn2us. No decommissioning, litigation, or adjudicated harm has been reported.

The case file concludes that a complaint channel cannot govern an error that produces no complaints. It names two defences: a standing second read of the letters the tool set aside, and telling letter writers there is a reading to question.

It sees the system's safety resting on one unpublished number: how often a caseworker pulls the original letter. That covers both acting on a flag and trusting that no flag means no problem.

What the design does well

The case file also credits the design. The tool is closed and encrypted, has no internet connectivity, redacts personal data after scanning, and makes no entitlement decision. So the data-leak risks that dominate other cases are low here. That leaves the unmeasured miss, and the unmeasured habit of checking the letter, as the questions that matter.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A pathway counts as closed once it passes on only a few mistakes. When a tool closes one, this text says mistakes stop passing along it. A network is self-correcting when its checks contain mistakes instead of letting them spread.

Six links are failure pathways at the start. Four involve the scanner: Scanner reads scanned post, Shortlist sent to caseworkers, Routing theme sent to mail staff, and Flag entered on the shortlist. The other two are Shortlist read as worklist and Ordinary queue read as worklist.

This case has a budget of 10 units. Each tool costs the same at every target level. Explore (No Targets) sets no targets.

Under Service Targets Only, the targets can be met within the budget. With no tool in use, they are not met. The cheapest ways cost 2 units and use one tool each: Escalate checks, Peer sharing rules, or Mark AI-written records. Most other combinations within the budget meet them too.

Under Service and Safety Targets, the targets also include closing every failure pathway. They can be met. The cheapest ways cost 7 units and use three tools: Escalate checks, Mark AI-written records, and either Store less data or Gate record entries.

Escalate checks closes the two pathways from the scanner to staff. Mark AI-written records closes Scanner reads scanned post and both pathways where staff read the queue. Store less data or Gate record entries closes Flag entered on the shortlist.

Under All Governance Targets, the targets can be met the same two ways, at the same cost. At both of these levels, every combination that meets the targets includes Escalate checks and Mark AI-written records. Each also includes Store less data, Gate record entries, or both.

None of the cheapest ways adds either check the case file argues for. Peer sharing rules adds the Second read of unflagged letters. Check with a second model adds the Independent accuracy check. Neither closes a failure pathway.

More checking is not always better here. Even if the budget allowed it, all ten tools at their standard settings cost 25 units, and 38 at their strongest. Either way they close every failure pathway, yet meet the targets at none of the three levels that set them. The added checks leave the scanner adding too little to the work.

Stylized model of a documented deploymentCaseworker documentation & copilots

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Whitemail-scanner-class upstream correspondence triage network: 6 components and 15 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 7 assumptions
  • assumed

    This network follows the pattern the case file documents for DWP's Whitemail Insights and Vulnerability Scanner. It is not a reconstruction of the actual tool. It is kept distinct from the Lab's other caseworker tools. Magic Notes writes social care records behind one review step that can drift. Nava and the Imagine LA Benefit Navigator draft cited answers that staff check before use. The cross-government Copilot experiment is an office assistant, whose lesson is that measuring a tool is not controlling it. Minute, a government-built meeting scribe piloted with councils, has a strong published governance record but no published accuracy evaluation. This tool instead reads incoming post first and acts only on what it flags.

  • baseline

    The network centres on a one-sided check around an error no one sees. A flagged case gets a caseworker who assesses it and can retrieve the original letter to verify it. A letter the scanner does not flag stays in the ordinary queue, with no second read for vulnerability. So staff judge the cases the tool surfaces, and no one judges the ones it drops. The sources describe neither a standing second read of the unflagged letters nor an independent accuracy check. The network includes both as checks that tools can add.

  • baseline

    The idea of an invisible error rests partly on the sources and partly on reading them. Two points are documented. Letter writers are not told the tool reads their letters: the impact assessment reported by the Guardian said they "do not need to know". Turn2us said, in the Guardian, that prioritising some cases inevitably deprioritises others. The rest is this Lab's reading of that design: a missed flag brings no complaint, and shows up, if at all, as harm blamed on something else. No source states that, and no such harm has been adjudicated.

  • assumed

    The check the system's safety rests on is not measured. The unique document identifier lets a caseworker retrieve the original letter to verify a flag. How often anyone does so is unpublished. So is how often staff disagree with or overrule a flag, and any estimate of vulnerable letters missed. How often the scanner errs is likewise assumed here, not measured. The record names three accuracy measures but discloses no values. No independent accuracy evaluation has been published.

  • assumed

    The network has no pathway for data leaving DWP, and that reflects the documented design. The system is hosted with end-to-end encryption and no internet connectivity. Personal data is automatically redacted after scanning. There is no commercial data-processing agreement with an outside processor: the vendor built the tool, and DWP keeps the intellectual property. The main risk here is the unseen miss, not a data leak. That leaves accuracy and correction as the questions that matter.

  • assumed

    Some sideways links in this network pass mistakes along, and one catches them. Caseworkers pass working practice, including habits formed around the flag thresholds, from desk to desk. One central scanner reads every letter, so any gap in what it recognises repeats across all of them at once. That is an assumption, not a measurement. The daily manual review is a real team check. The sources describe neither the second read nor the accuracy check. Outside scrutiny has come from press reporting, Turn2us, an independent analyst, and Parliament. No independent accuracy evaluation has been published.

  • assumed

    The potentially vulnerable claimants a missed flag would leave in the ordinary queue are not part of this network. What a miss means for the person who wrote the letter is documented in the case file, outside this network. The sources publish no measure of who a miss affects, and the Equality Analysis exists but is unpublished. This Lab models how errors move between the scanner, the staff, and the work queue. Every performance and time-saving figure cited for the tool is a claim by DWP or government, not a measurement.

What this example does not show

Show all 3 limitations
  • This example does not show the potentially vulnerable claimants a missed flag would leave in the ordinary queue. It shows how errors move between the scanner, the staff, and the work queue. It has no demographics and estimates no difference in harm between the people served. An Equality Analysis exists but is unpublished. The sources publish no measure of who a miss affects.
  • This example does not use accuracy figures, because none are published for this tool. The transparency record names three accuracy measures but discloses no values. No independent accuracy evaluation has been published. Also unpublished are how often staff retrieve the original letter, how often they overrule a flag, and any estimate of vulnerable letters missed. The invisible-error reading is this Lab's analysis of the documented design: claimants are not told, and staff work from a shortlist. Press criticism of the tool is reported concern, not adjudicated harm.
  • This example does not treat the throughput and time-saving figures as measurements. Each is a claim by DWP or its ministers and officials, with no independent verification. Daily volume is reported as 22,000 at the end of 2023 and in March 2024, and about 25,000 by November 2025. The fall from four to six weeks to 75% answered the same day comes from a trade press profile of DWP's Chief Digital and Information Officer. Earlier research based on freedom-of-information requests suspected a different supplier, Agilysis, than the Accenture (UK) Limited the transparency record names. How often the scanner errs is an assumption here, not a calibrated value.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • According to its Algorithmic Transparency Recording Standard record, published on November 27, 2025, the UK Department for Work and Pensions runs a Whitemail Insights and Vulnerability Scanner that reads roughly 25,000 scanned documents a day (reported as around 22,000 a day at end-2023 and in a March 2024 operator interview). Each document is passed first through the Vulnerability Scanner, a pre-trained open-source transformer doing zero-shot classification, which flags potentially vulnerable customers against eight prescribed themes including suicide and self-harm, domestic violence and abuse, and financial hardship; only documents not flagged as indicating vulnerability are relayed to Whitemail Insights for routing across nine themes. The output to trained staff is an anonymised daily report of flagged customers, and DWP states the tool does not make or influence benefit entitlement decisions. The record names precision, recall, and F1-score as its evaluation metrics but discloses no values, and no independent accuracy evaluation has been published.

    empirical
    • Government Department for Work and Pensions, Algorithmic Transparency Record: Whitemail Insights and Vulnerability Scanner (GOV.UK, 2025) https://www.gov.uk/algorithmic-transparency-records/whitemail-insights-and-vulnerability-scanner
    • Trade press Trendall, DWP taps AI to scan 25,000 letters a day and identify vulnerable citizens (PublicTechnology, 2025) https://www.publictechnology.net/2025/12/08/society-and-welfare/dwp-taps-ai-to-scan-25000-letters-a-day-and-identify-vulnerable-citizens/
    • Government UK Parliament Work and Pensions Committee, DWP use of artificial intelligence: correspondence (2023) https://committees.parliament.uk/publications/42458/documents/211057/default/
    • Trade press Corbridge (interview), How DWP is getting AI to work (Computing, 2024) https://www.computing.co.uk/interview/4188076/dwp-getting-ai
  • Guardian FOI reporting in January 2025 recorded that benefit claimants are not told the AI reads their correspondence: the internal data protection impact assessment stated that letter writers do not need to know about their involvement in the initiative, and the tool had been piloted since at least 2023 without appearing on the central government AI transparency register despite a ministerial mandate. The correspondence it processes can include national insurance numbers, health information, bank details, and children's details. Turn2us policy manager Meagan Levin voiced serious concerns, noting that prioritising some cases inevitably deprioritises others, so it is vital to understand how these decisions are made and ensure they are fair. The further reading that a missed flag on the unflagged residual therefore has no complaint channel and surfaces only as downstream harm is an analytical inference from the documented non-notification and shortlist design, not an adjudicated harm.

    empirical
    • Investigative Booth, Serious concerns about DWP use of AI to read correspondence from benefit claimants (The Guardian via inkl, 2025) https://www.inkl.com/news/serious-concerns-about-dwp-s-use-of-ai-to-read-correspondence-from-benefit-claimants
    • Trade press Toth, AI use for welfare system in doubt as scale of DWP setbacks revealed (The Independent via Yahoo News, 2025) https://www.yahoo.com/news/ai-welfare-system-doubt-scale-170440585.html
    • Advocacy Dent, Digital Welfare State edition 006 (ABD Consultancy, 2025) https://www.abdconsultancy.co.uk/blog/digitalwelfarestateedition006

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Caseworker documentation & copilots domain page.

Levers available here and the patterns behind them

Documented case histories