PAN Lab example
Justice Transcribe
The note that scores you: a copilot at the head of a risk pipeline
Justice Transcribe drafts summaries of probation meetings for the case record in England and Wales. The ministry's reoffending-risk tools draw on probation records too.
See more
Justice Transcribe is an AI tool the Ministry of Justice built in-house for probation staff in England and Wales. It records supervision meetings with people on probation, transcribes them, and drafts structured summaries for the case record. It makes no decisions itself.
How it is used
An officer records a supervision meeting with a person on probation. Supervision meetings are the regular meetings an officer holds with each person they supervise. Justice Transcribe transcribes the meeting, summarises it, and structures the summary for the case record. The officer reviews and accepts each summary, and owns the record.
Before the tool, officers copied handwritten notes into digital systems by hand. The ministry says the tool supports admin work. Its AI Action Plan for Justice says AI should "support, not substitute, human judgment."
The ministry's Justice AI Unit built the tool. It counts it as a benefit that summaries stay consistent as a person moves between officers.
How fast it spread
Justice Transcribe was piloted in Kent, Surrey, Sussex, and Wales. The AI Action Plan for Justice, published 31 July 2025, reported that the pilot cut note-taking time by 50 per cent. Officers rated it 4.5 out of 5.
Official transparency data give a national start date of 7 October 2025. On 23 October 2025 the government announced that more than 1,000 probation officers would be equipped with the tool. The same press release announced OpenAI's expansion into UK data hosting through the ministry's partnership.
More than 150,000 meetings had been summarised by 12 February 2026, and more than 800,000 by 2 June 2026. So meeting volume roughly quadrupled in the second four months, as the tool spread from about 1,000 officers to the whole service.
On 9 June 2026, at London Tech Week, the government said every probation officer in England and Wales had the tool. That was about a year after the pilot.
What the time savings rest on
Every time-saving figure is the ministry's own, or users' own reports. None has been measured independently.
The October 2025 announcement said AI technology across the ministry was expected to save up to 240,000 days a year. The published savings for Justice Transcribe rest on an operating assumption of about ten minutes saved per meeting. Probation Workforce Transformation, which the sources list among the bodies running the deployment, set that assumption. The department calls it illustrative and "not a formal statistical estimate." It does not account for differences in meeting length, type, practice, or summary quality.
On that assumption, the first transparency report gave an indicative 25,000 hours. The second gave about 133,333 hours over 800,000 meetings. It also gave a modelled forecast of about 450,000 hours a year, which is not a measure of time actually saved. The June 2026 announcement said Justice Transcribe alone could free up about 18,750 calendar days a year.
More than 63,000 voluntary reviews in the app average 4.7 out of 5. The ministry cautions that they "should not be interpreted as representative of all users."
What has not been published
No accuracy rate for Justice Transcribe's transcripts or summaries has been published. No rate of officer corrections or overrides has been published either. Nor has any audit of the quality of the records, or any independent evaluation of the tool.
The scale is counted to the meeting. Whether the summaries are right is not measured at all. The tool reached every probation officer before any of those figures existed.
The risk tools that draw on probation records
The same ministry runs algorithmic reoffending-risk assessment over probation records at high volume. Statewatch, a non-governmental organisation, reported in April 2025 that OASys-based risk prediction made more than 1,300 assessments a day, of prisoners and people on probation. That was 9,420 assessments in one week of January 2025. It draws on probation and prison caseload systems and the Police National Computer.
The ministry's own validation found that these predictions of reoffending were less accurate for some ethnic groups. For non-violent reoffending, that was all Black, Asian and Minority Ethnic groups. For violent reoffending, it was people of Black and Mixed ethnicity. That is a property of the risk tools, not of Justice Transcribe.
The ministry is rolling out a successor, Assess Risks, Needs and Strengths, during 2026. The AI Action Plan also sets Justice Transcribe beside a single offender identity system and AI-powered search of case materials for risk factors.
So the records Justice Transcribe helps write are among the records that high-volume risk scoring draws on. No published source documents a named data pipeline from the summaries into the risk tools. The documented connection is the shared case record.
No one has reported the tool being withdrawn, a lawsuit over it, or a ruling that it caused harm.
What outside commentary says
A peer-reviewed Probation Journal editorial by Jake Phillips, published online in December 2025, calls transcription tools the "first wave" of AI in probation. Citing policing research, it warns that "perceived efficiency gains have not materialised in practice."
The editorial flags bias and the erosion of professional judgment. It also flags model sycophancy: an AI's tendency to agree with what its user seems to want. And it raises an open question about who is accountable when an officer acts on a record an algorithm has shaped.
Mike Nellis wrote a critique for the Centre for Crime and Justice Studies, dated 26 March 2026. It records a concern of the Public Accounts Committee and the National Audit Office. They warned that the pace of new digital tools could "disrupt services, contribute to poor outcomes and staff stress."
The critique notes that "the MoJ does not have a strong history of implementing digital change programmes well." The MoJ is the Ministry of Justice. It also flags the OpenAI contract as a risk of pressure from a supplier.
What this case asks
The official account is reassuring. The tool is admin support, it decides nothing, and a person reviews and owns every record. The case file points to what that account leaves out: who else reads the record.
Officers read it to supervise people. The ministry's risk tools draw on probation records to score reoffending risk. The network assumes that includes the records Justice Transcribe helps write. A summary written to save ten minutes can become an input to a risk score. The sources describe no check between the two.
The case file also stresses the order of events. The tool reached the whole workforce before anyone published how often it is right. It argues for proving the checks before the next group of officers gets the tool, not after.
When a record serves two readers, an officer owning the record is not the same as the record being safe for a machine to read.
What this network is drawn from
This network follows the pattern the case file describes. It is not a reconstruction of the actual tool. It shows Justice Transcribe, the recordings, the officers, the case record, the OASys risk tools, the ministry's AI governance, and the recall and court-report step. People on probation are outside the network.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 12 units. Understand the system, the tool that pays for ongoing study of what the deployment is doing, costs 3 units under Service Targets Only. It costs 4 under the two higher levels. While it is in use, four tools cost 1 unit less each: Review on schedule, Escalate checks, Keep prompts neutral, and Upgrade model. Its stronger setting costs 6 units and takes 2 units off each, to no less than 1.
The network starts at a tipping point: its mistakes could either die out or build on each other. In a self-correcting network, mistakes die out before they build. Seven failure pathways are open at the start. They are Recording to Justice Transcribe, Summary to the officer, Summary into the case record, and History read at next meeting. The other three are Case record to OASys, Risk score to the officer, and Case record to recall and reports.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets are not met at the start. Mark AI-written records alone meets them, for 2 units. More than a thousand different sets of tools within the budget meet them.
Service and Safety Targets and All Governance Targets ask you to close every failure pathway, among other targets. Both can be met within the budget.
Under Service and Safety Targets, 15 different sets of tools meet the targets, and the cheapest cost 10 units. Every one includes Vet connections and Escalate checks. Vet connections is the only tool on offer that closes Case record to recall and reports. It also closes Recording to Justice Transcribe and Case record to OASys. Escalate checks is the only one that closes Summary to the officer and Risk score to the officer.
Each of those sets also includes Gate record entries or Store less data, which close Summary into the case record. Each also includes Mark AI-written records or Understand the system, which close History read at next meeting.
Lingering effects is a Dynamics setting in which damage outlasts its cause. It is on by default, off under Service and Safety Targets, and always on under All Governance Targets. With it on, Vet connections changes nothing unless Understand the system is also in use.
So under All Governance Targets, two sets meet the targets. One is Understand the system, Escalate checks, Vet connections, and Store less data, for 11 units. The other adds Keep prompts neutral, for 12. Gate record entries in place of Store less data closes the same pathways, but Justice Transcribe then adds too little to the work for that level.
The three checks the network draws are needed for no target. Check with a second model adds the check between Justice Transcribe and OASys. Assign a challenger and Peer sharing rules add the accuracy check on summaries. None of them closes a failure pathway. No tool on offer adds the check of the case record against its source.
More is not better here. Using every tool on offer, each at its strongest setting, costs 39 units, more than three times the budget. It meets the targets at none of the three levels that set them, because Justice Transcribe then adds too little to the work.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Justice-Transcribe-class multi-hop lineage copilot network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This network follows the pattern the Justice Transcribe case file describes. It does not reconstruct the actual tool. It draws the link from Justice Transcribe's summaries to the risk tools through the shared case record. No source documents a named data pipeline from the tool's output into the risk tools' inputs.
- baseline
The network assumes two AI tools share one record. Justice Transcribe helps write the legally required supervision record, and an officer reviews and accepts each summary. The ministry's reoffending-risk tools, reported at more than 1,300 assessments a day, draw on the same records. Officers then act on the score. So a line Justice Transcribe writes can become an unmeasured input to a high-stakes risk score. The sources describe no check between the two.
- baseline
The network assumes the rollout ran ahead of evaluation. The tool went from pilot to every probation officer in England and Wales in about a year. Meeting volume roughly quadrupled between the two published transparency reports. No accuracy, error-rate, or officer-correction evaluation was published, and no independent evaluation exists. Ministry governance publishes usage and satisfaction, not accuracy. So the network draws a standing accuracy check that the sources do not describe.
- assumed
The network assumes officers pass summary habits to one another. The Justice AI Unit counts it as a benefit that summaries stay consistent as a person moves between officers. The network assumes one shared tool also makes the records' wording uniform across the workforce. It also draws three checks the sources do not describe: between the two tools, on summary accuracy, and on the record before a recall or court report. The only scrutiny the sources describe comes from Parliament, researchers, and civil society groups, not an accuracy audit of the tool.
- assumed
The ministry's own validation found that the downstream risk tools predict reoffending less well for some ethnic groups. For non-violent reoffending, that was all Black, Asian and Minority Ethnic groups. For violent reoffending, it was people of Black and Mixed ethnicity. That is a property of the risk tools, not of Justice Transcribe. People on probation are not in this network, and the case file records that finding.
- assumed
Every time-saving figure for this tool is self-reported by users or asserted by government, not measured independently. That covers a 50 per cent cut in note-taking time and an illustrative ten minutes per meeting. It also covers a modelled 450,000 hours a year, stated as about 18,750 calendar days. The ministry's figure of up to 240,000 days a year covers all its AI tools. No transcription-accuracy evaluation is published. The network shows how mistakes pass between tools, officers, and records, and estimates none of these figures.
What this example does not show
Show all 2 limitations
- This example does not show people on probation or what happens to them. In the ministry's own validation, the downstream risk tools predict reoffending less well for some ethnic groups, differently for each kind of reoffending. That is a property of the risk tools, not of Justice Transcribe. The network includes no demographics and estimates no difference in harm. The case file records that finding.
- This example does not show a documented data pipeline. No source documents one from Justice Transcribe's output into the risk tools' inputs, so the link is drawn through the shared case record. Every time-saving figure is self-reported or asserted by government, not measured independently. No transcription-accuracy evaluation is published. This example shows how the tool fits into probation work, and estimates none of those figures.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The Ministry of Justice built an in-house AI transcription and summarisation copilot, Justice Transcribe, for probation staff in England and Wales, scaling it from a pilot to more than 1,000 officers in October 2025 and to every probation officer by June 2026, with official transparency data recording more than 800,000 supervision meetings summarised between 7 October 2025 and 2 June 2026; the reported time-savings are the ministry's own and rest on an operating assumption the department itself labels illustrative, and no transcription-accuracy rate, officer correction rate, or independent evaluation of the tool has been published.
empirical- Government Justice AI Unit, Ministry of Justice, Justice Transcribe in Probation (2026) https://ai.justice.gov.uk/our-work/justice-transcribe
- Government Ministry of Justice, AI Action Plan for Justice (GOV.UK, 2025) https://www.gov.uk/government/publications/ai-action-plan-for-justice/ai-action-plan-for-justice
- Government Ministry of Justice and HM Prison and Probation Service, Justice Transcribe data 7 October 2025 to 2 June 2026 (transparency data, GOV.UK, 2026) https://assets.publishing.service.gov.uk/media/6a1eafe265bc5f798327f61f/Justice-transcribe-report-2-june-2026.pdf
- Government Ministry of Justice and DSIT, OpenAI to expand into UK data hosting after major growth deal (GOV.UK press release, 2025) https://www.gov.uk/government/news/openai-to-expand-into-uk-data-hosting-after-major-growth-deal
- Government Ministry of Justice, AI tech ambition to deliver smarter justice for victims (GOV.UK press release, 2026) https://www.gov.uk/government/news/ai-tech-ambition-to-deliver-smarter-justice-for-victims
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
- Vet connections — Connection authorization
- Escalate checks — State-feedback vigilance
- Keep prompts neutral — Framing and mirroring reduction
- Gate record entries — Human-in-the-loop write gating
- Mark AI-written records — Provenance labeling
- Understand the system — Understand the system
- Store less data — Data minimization
- Peer sharing rules — Peer-edge governance
- Upgrade model — Improve the model
Documented case histories
- Justice Transcribe
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check