PAN Lab example
Magic Notes (Beam)
The drafted record: a case-notes copilot
UK council social care staff draft case notes, the client's record, with Magic Notes, AI software from the company Beam. They review each draft first.
See more
Magic Notes is a generative AI tool from the company Beam, used by UK councils in adult social care. It records conversations such as assessments and visits, and drafts case notes and assessments from them. A practitioner, the social worker or other care worker who held the conversation, reviews each draft before it enters the case notes.
How it is used
The case file says vendor materials and council evaluations describe its adoption across a large number of local authorities. They claim time savings on assessment paperwork.
Somerset Council reported that its social workers save time on admin with it. The claimed figures are write-ups 48 to 65 percent faster, and about 11 to 12 hours saved a week. The sources read for this case do not say who measured them, or whose hours were saved.
Who checks the notes
The practitioner who held the conversation reviews each draft, can edit it, and accepts it before it enters the case notes. The sources list a team manager's review above that. They do not say whether a manager reviews a note before or after it enters the case notes. They list the council's data protection and practice governance function as the authority over how the tool is used.
The case file asks three questions the sources do not answer. What share of drafts are actually edited? Who audits accepted notes against the source audio? How are AI-drafted records labelled for later readers?
Why the record is what needs governing
Magic Notes makes no decision about anyone. That is what makes the risk easy to miss.
A wrong detail in a draft, or a made-up one, can survive a review cut short by time pressure and enter the permanent case record. Later readers and later decisions then take it as fact. The network assumes readers take it as the practitioner's own writing, and that later drafts draw on it.
The case file warns that under time pressure, review drifts toward approval. The sources do not say how closely practitioners read drafts in practice.
Why this case matters
The case file calls this an ongoing experiment, not a failure being examined afterwards. It is a live test of whether a human review step holds up at scale. The case file notes that the design keeps a human review step between the drafted note and the official record.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Here a mistake is, for example, a made-up detail in a draft that a practitioner accepts into the case notes. A pathway counts as closed once few mistakes pass along it. The notes themselves still move.
This case has a budget of 11 units. Explore (No Targets) sets no targets. There, one tool costing 2 units is enough to stop mistakes building on one another across the network. Mark AI-written records is one.
Under Service Targets Only, one tool also meets the targets. Five tools each do it alone, for 2 units: Mark AI-written records, Keep skills sharp, Peer sharing rules, Assign a challenger, and Escalate checks. At the default setting, Upgrade model and Gate record entries also do it alone, for 3 units.
Under Service and Safety Targets, the targets can be met, but only with the whole budget. Both higher levels ask you to close every failure pathway, among other targets. Four combinations of five tools meet them, each costing all 11 units. Each has Mark AI-written records, Keep prompts neutral, and Escalate checks. Each then adds Gate record entries or Store less data, and Assign a challenger or Peer sharing rules.
Under All Governance Targets, two of those combinations meet the targets: the two with Gate record entries. Lingering effects is a Lab setting that lets an effect outlast its cause. It is always on at that level, off at Service and Safety Targets, and on by default at the two lower levels. With it on, Vet connections, Store less data, and Peer sharing rules work at reduced strength. Full strength needs Understand the system, a tool that funds ongoing study of what the deployment is really doing. This case does not offer it, so Store less data falls short at All Governance Targets.
More is not better here. Every tool at its strongest setting at once closes every failure pathway, but costs 35 units, far over the budget. It also leaves Magic Notes no longer clearly helping the work, so it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Magic-Notes-class documentation copilot network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
This network follows the pattern of an AI tool that drafts case notes, as the Magic Notes (Beam) case file describes it. It is not a reconstruction of the product itself.
- assumed
The network assumes three peer effects. Practitioners pass drafting shortcuts to one another. One tool drafting every note makes the notes sound alike. Colleagues still check each other's drafts when time allows.
- assumed
The network assumes practitioners edit drafts less closely as time pressure rises. It starts from practitioners making some real edits. The sources do not measure how closely practitioners edit. The case file warns that under time pressure, review drifts toward approval.
- assumed
The network gives the team manager a review of their own, above the practitioner's review. The sources list a managerial review for this tool. They do not say how often it happens.
- assumed
The network assumes later readers take drafted text that a practitioner accepted as the practitioner's own writing.
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Gate record entries — Human-in-the-loop write gating
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Upgrade model — Improve the model
- Vet connections — Connection authorization
- Keep skills sharp — Deskilling-arrest mandate
- Store less data — Data minimization
- Peer sharing rules — Peer-edge governance
- Assign a challenger — Structured dissent
- Escalate checks — State-feedback vigilance
Documented case histories
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check