Skip to content

PAN Lab example

An ambient AI scribe at a multi-specialty health system

One number and two outcomes: an ambient scribe that split by clinician

Sutter Health's AI scribe served two clinician groups. Of primary-care physicians, 85.8 percent reported improved satisfaction, against 36.4 percent of specialists. One average hides that.

See more

The scribe is a commercial AI documentation platform built into Sutter Health's electronic health record, from a vendor the published reports do not name. It transcribes a patient's visit and drafts the clinical note. A clinician reviews, edits, and signs the draft before it becomes part of the permanent record.

Where it was used

Sutter Health is a large health system with many specialties in Northern and Central California. The evaluation compared two groups of clinicians who used it: primary-care physicians and medical specialists. Both groups used the same tool, in the same system, under the same workflow.

What the evaluation found

Stults and colleagues published a peer-reviewed evaluation of the scribe in JAMA Network Open in 2025. It measured what many deployment reports do not: how the benefit was spread across clinicians, and where it did not appear.

Note time per appointment fell from 6.2 to 5.3 minutes. Cognitive load, the mental effort the work takes, fell. It was measured with the NASA Task Load Index, a standard workload questionnaire.

Burnout went from 42.1 to 35.1 percent. That change was not statistically significant, so the study could not show it was more than chance.

Who the benefit reached

The benefit split sharply by clinician group. Of primary-care physicians, 85.8 percent reported improved satisfaction. Of medical specialists, 36.4 percent did.

The case file reads this as two different outcomes inside one headline number, not noise around a single average. A single benefit figure overstates the effect for the group the tool helps least. A report of "improved satisfaction" without the split reports the primary-care result as if it held for everyone.

So watch what a single benefit number hides: who the tool actually helps, and who carries the same review work without the relief it promised.

Why the burnout result matters

The case file calls this an honest limit for every AI scribe case in the domain: workload fell, but an effect on burnout was not shown.

The evaluation was being honest about what it could and could not show. Larger benefit figures from other scribe deployments should be read against it, not in place of it.

What the evaluation did not measure

It measured time, workload, burnout, and satisfaction. It did not measure how many errors the drafts contained.

The sources give the note-time figure for the evaluation as a whole, not for each group. They do not show Sutter repeating the measurement by group on a schedule.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other.

This case has a budget of 11 units. Explore (No Targets) sets no targets. Under each of the three other levels, the targets can be met within that budget.

Under Service Targets Only, the cheapest combinations cost 4 units. One uses two tools: Review on schedule and Escalate checks.

Service and Safety Targets and All Governance Targets also require every failure pathway to be closed. There, every combination that meets the targets uses at least four tools, and the cheapest ones cost 9 units. One is Mark AI-written records, Keep prompts neutral, Escalate checks, and Store less data.

Stylized model of a documented deploymentClinical documentation copilots (ambient scribes)

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Ambient-scribe-class with operator heterogeneity network: 5 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    The model assumes specialists spend more effort correcting the scribe's drafts than primary-care physicians do. It follows the evaluation's split: 85.8 percent of primary-care physicians reported improved satisfaction, against 36.4 percent of medical specialists. The case file reads that split as specialists carrying the same review work on a fluent draft without the same relief. No source measures correction work by group. The model also assumes the clinicians' review workload exceeds the time they have for it, reflecting that split rather than an even load. The evaluation reported satisfaction separately for each clinician group, which is how the split was found at all. The sources do not show Sutter tracking results by group on a schedule rather than reporting one average.

  • baseline

    This models the pattern of one scribe serving two clinician groups, as the Sutter case file documents it. It is not a reconstruction of Sutter's actual systems. The layout is the standard one for AI drafting into a permanent record. The scribe drafts, a clinician edits and signs, and the record keeps the signed note. What differs is that two groups of clinicians use the same scribe, and the measured benefit split sharply between them. So a single benefit figure overstates the effect for the group the tool helps least.

  • assumed

    The two clinician groups use the same tool and the same workflow. The documented difference is a measured outcome: 85.8 percent of primary-care physicians reported improved satisfaction, against 36.4 percent of medical specialists. The Lab does not compute that split. It is an observation recorded in the case file. The network carries it only as what a check measuring results by group would find.

  • baseline

    The evaluation sets an honest limit for scribes like this one. Note time and cognitive load, the mental effort the work takes, fell. But the change in burnout, from 42.1 to 35.1 percent, was not statistically significant. Larger benefit figures from other deployments should be read against this result, not in place of it. The benefit is spread unevenly across clinicians, not one number.

  • assumed

    No patient care outcome is part of this network. The patients whose visits are transcribed are outside it. The satisfaction split, the note-time figures, and the burnout result are recorded in the case file. None of them is computed from anything in this network.

What this example does not show

Show all 2 limitations
  • This example shows no patient care outcome. It follows how a mistake by the scribe, a clinician, or the note record can be repeated by the others. The patients whose visits are transcribed are outside it. The satisfaction split, the note-time figures, and the burnout result come from the case file. The Lab does not compute them.
  • The two clinician groups share the same tool and workflow. The split of 85.8 percent against 36.4 percent reporting improved satisfaction is a measured outcome, recorded outside this example. The Lab does not compute the benefit or how it divides between the groups.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note time per appointment reduced (6.2 to 5.3 minutes) and NASA-TLX cognitive load reduced — but the burnout change (42.1 to 35.1 percent) was not statistically significant, the domain's honest null bound of cognitive-load relief without a demonstrated burnout effect.

    empirical
    • Peer-reviewed Stults, C.D., Deng, S., Martinez, M.C., et al. (2025). Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians. JAMA Network Open, 8(5), e258614. https://doi.org/10.1001/jamanetworkopen.2025.8614 https://pubmed.ncbi.nlm.nih.gov/40314951/
  • In the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians reported improved satisfaction against 36.4 percent of medical specialists — the same tool, in the same system, under the same workflow, helping one operator class and largely failing another, so any uniform service term overstates the effect for the group it helps least.

    empirical
    • Peer-reviewed Stults, C.D., Deng, S., Martinez, M.C., et al. (2025). Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians. JAMA Network Open, 8(5), e258614. https://doi.org/10.1001/jamanetworkopen.2025.8614 https://pubmed.ncbi.nlm.nih.gov/40314951/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical documentation copilots (ambient scribes) domain page.

Levers available here and the patterns behind them

Documented case histories