Skip to content

PAN Lab example

Ambient scribe RCT + monitoring playbook

The trial and the playbook: an evidenced and monitored ambient scribe

UW Health tested Abridge, an AI that drafts clinical notes, in a randomized trial, and published a playbook for monitoring it in use.

See more

Abridge is a commercial ambient AI scribe. At UW Health, the University of Wisconsin Hospital and Clinics, it records a patient visit and drafts the clinical note in the Epic electronic health record. UW Health uses it under a software-as-a-service agreement with Abridge.

How it is used

The clinician reviews, edits, and signs each draft. Only then does it become part of the permanent medical record. Later clinicians read that record as fact.

What the trial found

UW Health tested Abridge in a pragmatic randomized controlled trial, published in NEJM AI in 2025. It ran for 24 weeks and included 66 practitioners and 71,487 notes. Of those notes, 38 percent were drafted by the AI.

Each practitioner was assigned at random a time to start using Abridge, a design called stepped-wedge. That gives a cause-and-effect estimate, not a figure from the deployer's own dashboard.

Work exhaustion fell significantly. The trial reported that it found no significant change in professional fulfillment. Practitioners spent roughly 22 minutes a day less on documentation. Diagnostic coding accuracy improved. Diagnostic codes are the standard codes for a patient's conditions that insurers use to pay for care.

Why the trial matters

Among the ambient scribe cases in this Lab, this trial has the strongest study design. Some other deployments in this domain report their benefits in large first-party reports, written by the deployers themselves. The trial's figures are what those numbers should be measured against.

UW Health's own team ran the trial, and it was peer reviewed. The public sources for this case show no standing outside audit beyond that peer review.

What the trial did not measure

The trial measured practitioner well-being and time. It did not measure how often a note contains an error. A separate 2025 evaluation of one ambient scribe found hallucinations, content the AI made up, in 31 percent of its notes, against 20 percent of physicians' own notes. That figure belongs to other cases in this domain, not to this trial.

The monitoring playbook

The same team published an open operations playbook for monitoring the safety and effectiveness of ambient AI in everyday use. It was built around this deployment.

In most deployments, the monitoring function is assumed rather than shown. Here it is a document that others can inspect, copy, and hold UW Health to.

The coding risk

The trial found that diagnostic coding accuracy improved. A 2025 policy brief in npj Digital Medicine names the risk around that gain: the coding arms race.

The brief argues that AI scribes raise the intensity of billing and risk-adjustment coding. Coding intensity is how many codes, and how severe, a note supports. Payers, the insurers who pay for care, respond by downgrading codes and adjusting risk scores. The escalation grows. The clinician who attests to a code the AI suggested carries the liability.

Better coding can shade into pressure to code higher than the care supports. The public sources for this case call the playbook UW Health's stated answer to exactly that risk. The public sources for this case do not show the arms race happening at UW Health.

What to watch

Many scribe cases show a benefit. Two things here deserve closer watching. One is what the trial did and did not measure. The other is whether the monitoring stays aimed at the coding risk that better documentation creates.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest way costs 2 of this case's 11 budget units and uses a single lever.

Under Service and Safety Targets and under All Governance Targets, the targets can also be met within the budget. Both levels ask you to close every failure pathway, among other targets. The cheapest combination costs 9 of the 11 units. No combination of fewer than four levers meets them.

Stylized model of a documented deploymentClinical documentation copilots (ambient scribes)

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Ambient-scribe-class with a published monitoring playbook network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    The network assumes UW Health's monitoring takes Abridge's drafts as they are made. That comes from UW Health's published playbook for monitoring ambient AI in everyday use. It is what separates this case from the Kaiser Permanente scribe case, whose quality program samples signed notes. The sources do not say whether it samples drafts or signed notes. The network also assumes clinicians have more capacity to review drafts than the drafts demand. The evidence is a 24-week randomized trial of 66 practitioners, a small group with operational support, not a large rollout. The network includes the trial as a check on the tool because a pragmatic randomized controlled trial is the strongest cause-and-effect evidence in this domain.

  • baseline

    This network follows the ambient scribe pattern: the AI drafts a note that becomes part of a permanent record. It is not a reconstruction of UW Health's actual system. It has the same shape as the Kaiser Permanente scribe case. What sets it apart is its evidence and its monitoring, not its shape. The evidence is a stepped-wedge, individually randomized trial of 66 practitioners and 71,487 notes. The monitoring is a published playbook for monitoring the scribe in everyday use.

  • baseline

    The network includes the randomized trial as a check on the tool's benefit. The trial found work exhaustion significantly reduced and about 22 minutes a day of documentation time returned. It reported that it found no significant change in professional fulfillment. That makes it the strongest cause-and-effect evidence among the ambient scribe cases, and an honest one. The trial measured well-being and time, not how often notes contain errors. A separate 2025 evaluation of one ambient scribe found hallucinations, content the AI made up, in 31 percent of its notes, against 20 percent of physicians' own notes. That finding belongs to other cases in this domain, not this one.

  • assumed

    The network assumes UW Health's monitoring is real, because it exists as a published operations playbook rather than an assumed practice. That is rare: here oversight is a document anyone can cite and inspect. The monitoring has to stay aimed at the documented coding arms race. Better AI documentation raises coding intensity, payers respond, and the liability of clinicians who attest to the codes grows.

  • assumed

    The network does not include patients' care outcomes. It covers how mistakes pass between the scribe, the clinicians, the record, and the monitoring. The patients whose visits are recorded are outside the network. The trial's estimates, the monitoring playbook, and the coding risk are described in the case file. Nothing in this network computes them.

What this example does not show

Show all 2 limitations
  • This example does not show whether patients' care got better or worse. It shows how mistakes pass between the scribe, the clinicians, the record, and the monitoring. The patients whose visits are recorded are outside the network. The trial's estimates, the monitoring playbook, and the coding risk are described in the case file, not computed here.
  • The trial measured practitioners' well-being and time, not how often notes contain errors. A separate 2025 evaluation of one ambient scribe found hallucinations, content the AI made up, in 31 percent of its notes, against 20 percent of physicians' own notes. This Lab treats that figure as typical of such scribes and carries it in other cases. It is not a finding of this trial.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized trial of an ambient scribe across 66 practitioners and 71,487 notes (38 percent AI-generated), which found work exhaustion significantly reduced, professional fulfillment unchanged (a recorded null), roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the larger first-party deployment reports, this is a randomized estimate of the well-being and time effects — though it measured practitioner well-being and time, not per-note error rates.

    empirical
    • Peer-reviewed Afshar, M., Baumann, M.R., Resnik, F., et al. (2025). A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being. NEJM AI. https://doi.org/10.1056/AIoa2500945 https://pubmed.ncbi.nlm.nih.gov/41625485/
  • The same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monitoring of ambient AI in production — the rare case where the organization-side monitoring function exists as a citable, designed subsystem rather than an assumed practice. That monitoring is the org's stated answer to a documented system-level risk of the technology: a coding arms race, in which better AI documentation raises coding intensity, payers recalibrate in response, and clinician attestation liability grows — so the improved coding accuracy the trial measured sits next door to an upcoding pressure the monitoring is meant to watch.

    empirical
    • Academic Afshar, M., et al. (2025). A Novel Playbook for Pragmatic Trial Operations to Monitor and Evaluate Ambient Artificial Intelligence in Clinical Practice. NEJM AI. https://doi.org/10.1056/AIdbp2401267 https://ai.nejm.org/doi/full/10.1056/AIdbp2401267
    • Peer-reviewed Dai, T., Kvedar, J.C., & Polsky, D. (2025). Policy brief: ambient AI scribes and the coding arms race. npj Digital Medicine. https://doi.org/10.1038/s41746-025-02272-z https://www.nature.com/articles/s41746-025-02272-z

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical documentation copilots (ambient scribes) domain page.

Levers available here and the patterns behind them

Documented case histories