PAN Lab example
Kaiser Permanente ambient AI scribe
The draft becomes the record: a well-governed ambient scribe
Kaiser Permanente physicians used an AI scribe that drafts visit notes. The physician edits and signs each draft, which then enters the medical record.
See more
Kaiser Permanente's ambient AI scribe is commercial software that listens to a visit and drafts the clinical note, which the clinician then edits and signs. The Permanente Medical Group ran it in Northern California. The published papers do not name its maker.
Who ran it and where
The Permanente Medical Group is the medical group of Kaiser Permanente Northern California. It ran a 10-week pilot of the scribe, then scaled it up.
From October 2023 to December 2024, 7,260 of its physicians used the scribe in 2,576,627 patient visits. It is the largest documented ambient-scribe deployment.
What Kaiser reports
Kaiser reports about 16,000 hours of documentation time saved, and sustained support for the scribe from its physicians. Kaiser published the pilot results in 2024 and the scale results in 2025, both in the journal NEJM Catalyst.
These are Kaiser's own measurements. They were peer-reviewed, but no one independent of Kaiser made them.
How each note is checked
The deployment runs two checks on the scribe's notes. The clinician reviews, edits, and signs every note before it is stored. A standing internal quality-assurance program also reviews samples of the scribe's notes across the deployment.
The program is a designed function with a real cost, not a default setting. The sources read for this case do not say how many notes it samples, or what it compares them against.
Why the record is what needs governing
A drafted note does not stay a draft. Once the clinician signs it, later clinicians read it as fact, and later tools take it in as training or context.
So a mistake that survives the review can be copied forward into later care. The clinician's review and the quality-assurance program guard the permanent record, not a passing alert.
What independent studies found
A 2026 study in JAMA followed 8,581 clinicians at five other health systems, 1,809 of whom used scribes. It found a more modest effect. Clinicians who used scribes spent about 13 fewer minutes a day in the health record and 16 fewer on documentation. Their after-hours work did not fall meaningfully.
A 2025 study of one ambient scribe found hallucinations, meaning made-up content, in about 31% of its notes under structured review. It found the same kind of error in about 20% of notes physicians wrote. The ambient notes were more thorough but less accurate.
Neither study measured Kaiser's scribe. Kaiser's figures describe its own system, not a guarantee for every scribe of this kind.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.
Explore (No Targets) sets no targets. Under Service Targets Only, one tool is enough to meet the targets. The cheapest is Mark AI-written records, at 2 of this case's 11 budget units.
Under Service and Safety Targets and All Governance Targets, the targets can also be met. Both levels ask you to close every failure pathway, among other targets. The cheapest combinations cost 9 of the 11 units and use four tools. One such set of four is: Mark AI-written records, Keep prompts neutral, Escalate checks, and Store less data.
More is not better here. Pulling every tool at its strongest setting at once closes every failure pathway but costs far more than the budget. It also leaves the scribe no longer clearly helping the work, so it misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Ambient-scribe-class generation into the record network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
Among the deployments reviewed for this case, Kaiser's is the only one documenting two things. One is a record from a small trial through full rollout. The other is an internal quality-assurance program over AI output. So this example draws that program taking samples of notes from the record. It assumes a heavy review workload, because about 2.58 million visits each produced a draft that someone had to read. The program samples notes rather than reviewing every one. No study of Kaiser's own time-saved figures by anyone independent of Kaiser appears in the record.
- baseline
This example follows the pattern the ambient-scribe case file documents. It does not rebuild Kaiser's actual system. The AI drafts a note that is stored in a permanent record and read later as fact. So what needs governing is the entry into the record, not an alert. That is why this case is kept apart from clinical alert tools.
- baseline
Two pathways carry the risk of copying mistakes forward. Earlier notes give the scribe context for its next draft, and later clinicians and tools read earlier notes as established fact. A mistake nobody catches can then last a long time as inherited fact. One study found hallucinations, meaning made-up content, in about 31% of the notes an ambient scribe drafted, under structured review. So the clinician's review and the quality-assurance program guard the record itself. A review reduced to a rubber stamp lets mistakes into it.
- assumed
This case shows a well-governed deployment. The clinician reviews, edits, and signs every note before it is stored. A standing quality-assurance program also reviews samples of the scribe's notes across the deployment. It is a real function with a real cost. This example assumes the program runs as a standing function.
- baseline
The independent benefit check stands for the limit on Kaiser's own figures. Kaiser's benefit numbers are its own measurements. An independent study of 8,581 clinicians at five health systems found a more modest effect, with no meaningful after-hours relief. Kaiser's scale figures describe its own system, not a guarantee for every scribe of this kind.
- assumed
This example draws no care outcome. It shows how mistakes move between the scribe, the clinicians, and the records. The patients whose visits are recorded are not drawn. Coding intensity means recording more billable diagnoses. The time-saved figures, the hallucination rates, and the risk of coding intensity come from the case file. Nothing in this diagram computes them.
What this example does not show
Show all 2 limitations
- This example shows no care outcome. It shows how mistakes move between the scribe, the clinicians, and the records, not what happens to patients. The patients whose visits are recorded are not drawn. Coding intensity means recording more billable diagnoses. The time-saved figures, the hallucination rates, and the risk of coding intensity come from the case file. Nothing in this diagram computes these.
- Kaiser's benefit numbers are its own measurements, peer-reviewed but not independent. An independent study across five other health systems found a more modest effect. The 31% hallucination figure comes from a study of one ambient scribe under structured review. It is not a measurement of Kaiser's system.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7,260 physicians and 2,576,627 patient encounters over fourteen months, with roughly 16,000 hours of documentation time saved and sustained physician support measured along the way. The system records the visit and drafts the clinical note; the clinician edits and signs, and the model-to-record write is gated both by that clinician review and by a standing internal quality-assurance program over the AI output — a real subsystem with a real cost, because the drafted note becomes a permanent record that later clinicians and later tools read as fact.
empirical- Academic Tierney, A.A., Gayre, G., Hoberman, B., et al. (2024). Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst Innovations in Care Delivery. https://doi.org/10.1056/CAT.23.0404 https://catalyst.nejm.org/doi/full/10.1056/CAT.23.0404
- Academic Tierney, A.A., et al. (2025). Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. https://doi.org/10.1056/CAT.25.0040 https://divisionofresearch.kaiserpermanente.org/ai-assisted-notetaking-gains-steady-support-from-kaiser-permanente-physicians/
The scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and should be read as that system's dashboard rather than a guarantee of the product class: a multisite study of 8,581 clinicians across five health systems found more modest effects — on the order of 13 to 16 fewer minutes per day with no meaningful after-hours relief — and a validated per-note evaluation found hallucinations in about 31 percent of ambient-generated notes under structured review, versus about 20 percent of physician-written gold-standard notes, making ambient notes more thorough but less accurate. The clinician review and quality-assurance program are the controls that stand between that error rate and a contaminated permanent record.
empirical- Peer-reviewed Rotenstein, L.S., et al. (2026). Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence-Powered Scribes: A Multisite Study. JAMA. https://doi.org/10.1001/jama.2026.2253 https://pubmed.ncbi.nlm.nih.gov/41920565/
- Peer-reviewed Palm, K.H., Manikantan, K., Mahal, N., Belwadi, S.K., & Pepin, R.J. (2025). Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence, 8. https://doi.org/10.3389/frai.2025.1691499 https://pmc.ncbi.nlm.nih.gov/articles/PMC12586549/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
All of them in context on the Clinical documentation copilots (ambient scribes) domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
- Upgrade model — Improve the model