Skip to content

PAN Lab example

Epic Sepsis Model

Switched on before anyone checked: a proprietary sepsis model at scale

Hundreds of hospitals switched on the Epic Sepsis Model before anyone independent checked it. When researchers did, it caught about a third of sepsis cases.

See more

The Epic Sepsis Model is proprietary software that Epic built into its electronic health record system. It scores hospitalized patients continuously and alerts clinicians to those it predicts may have sepsis. Hundreds of US hospitals switched it on while Epic kept the model closed to outside inspection, which STAT News called a corporate firewall.

How it was used

Epic sells the electronic health record, and the model came built into it. Each hospital that ran it switched it on and set its own alert thresholds and routing. Hundreds of hospitals across the United States did so.

The model belongs to Epic. The hospitals running it could not readily inspect what it did or how well it worked.

What the 2021 validation found

Researchers at Michigan Medicine, one of the health systems running the model, tested it on 38,455 of their own hospitalizations. JAMA Internal Medicine published the study in 2021. Unlike the developer-led evaluations of other tools in this domain, it was independent of the company that built the model.

The model caught about 33 percent of sepsis cases. It flagged 18 percent of all hospitalizations. Its positive predictive value was near 12 percent: roughly one in eight of its sepsis predictions was right.

Its area under the curve was 0.63. That measure shows how well a score ranks a real case above a non-case, where 0.5 is a coin flip.

STAT News put the alert burden at roughly 109 alerts for every true case of sepsis. The sources read for this case do not explain how that figure was counted, or how it relates to the 12 percent figure.

What the investigation reported

STAT News investigated the model in 2021. It reported that Epic had not fully examined the model's real-world performance before selling it.

It also reported that the model used undisclosed inputs (what data scientists call features), such as antibiotic-order data. Those let the model partly predict a treatment clinicians had already started, rather than the condition itself. They explain part of the gap between Epic's strong internal numbers and the weak external ones.

Two failures that compound

The first is alert fatigue. At roughly 109 alerts per true case, by STAT News's count, most alerts are false. A clinician interrupted that often learns to dismiss the alert reflexively, and the alert becomes noise the workflow routes around.

The second is vendor opacity. Epic kept the model closed to outside inspection, which made independent scrutiny difficult. The hospitals switching the model on could not readily inspect what it did or how well it worked. The validation that exposed the gap came from researchers at one of those hospitals, not from Epic.

What changed after the criticism

After the outside criticism, Epic overhauled the model, as STAT News reported in October 2022. It retrained the model, moving to training on each hospital's own data. It changed its definition of when sepsis begins, and reduced its reliance on antibiotic-order inputs.

In 2026, JAMA Network Open published a validation of the updated model at four health systems. It followed 227,091 patient encounters forward in time. Its area under the curve was 0.82 to 0.92. Its positive predictive value was 13 to 26 percent, and results varied substantially between sites.

The authors urged local validation and alert-silencing strategies, rather than trusting the model out of the box. The model got better, but only after independent scrutiny forced the issue. By then it had run at scale for years.

What this case asks

This case asks about the governance that should have come first. It asks for an independent check before the model runs at hundreds of sites. It asks for control of the alert volume. It asks for terms that let a hospital see what its vendor ships.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest combination of tools that meets them uses two tools and costs 4 of this case's 12 budget units.

Under Service and Safety Targets and under All Governance Targets, the targets include closing every failure pathway. They can be met, but only just. Every combination of tools was checked. Exactly one meets them: four tools, costing all 12 units.

This case offers no tool that changes the model itself. The model belongs to Epic, so the hospital running it cannot redesign it. The control the hospital holds over the model itself is its terms with Epic.

Stylized model of a documented deploymentClinical decision support & deterioration alerting

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Proprietary-EHR-sepsis-class model at scale network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 7 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 8 assumptions
  • assumed

    This example has the same parts as the Lab's example of TREWS, another sepsis alert system. Both records describe one arrangement: a model inside the electronic health record scores a continuous stream of patient data and alerts the treating clinician. Both also have an external validation check. Neither record documents a part the other lacks, so adding one would be invention. The two examples differ in how their parts behave, following their sources. The live patient-data stream is its own input here, because the record treats it as the model's defining input. Here that data also included antibiotic orders as clinicians entered them, so the score partly reflected decisions already made.

  • assumed

    This example treats the model's reading of the health record as a major source of error. STAT News reported that the model used undisclosed inputs, including antibiotic-order data. So it partly read the treatment decision it was meant to predict. That is the sharpest documented case in this Lab of a model learning from the record it scores. It is documented, not inferred: the model learned from treatment decisions clinicians had already made.

  • assumed

    This example treats one model at hundreds of hospitals as a major source of repeated error. One vendor built and updates the proprietary model, and hundreds of hospitals run it. Nothing in this Lab is more uniform across sites by design. A blind spot in it is the same blind spot everywhere it runs.

  • assumed

    This example treats the check named External validation of the model as a check that happened late, not as one that never happened. That sets this case apart from every case whose failure is an audit that never took place. An external validation over 38,455 hospitalizations was done and published in 2021. It was damning: poor ranking of patients, a third of cases caught, and, by STAT News's count, about 109 alerts per true case. The model still stayed in use at scale until Epic chose to overhaul it. So what this case lacks is not the check but the authority to act on one. A finding that no one with power to stop a deployment acts on is not oversight.

  • baseline

    This example follows the failure the case file documents: a model switched on at scale before anyone independent validated it. It is not a copy of Epic's actual model. It opens already overloaded, as the case documents, with heavy alert volume and clinicians checking less. A proprietary model was switched on across hundreds of hospitals before independent validation. It caught about a third of sepsis, at roughly 109 alerts per true case by STAT News's count.

  • assumed

    This example assumes the alerts come in very large numbers and the clinicians' checking is worn down. At roughly 109 alerts per true case, by STAT News's count, most alerts are false. So the clinicians' review turns into reflexive dismissal. That is alert fatigue: a tool meant to help becomes noise the workflow routes around. The case file names aggressive tiering and alert-silencing as the correction. The authors of the 2026 validation urged alert-silencing strategies.

  • assumed

    The defining gap is the check named Local validation and review: no independent validation ran before the model was switched on at scale. Epic kept it closed to the outside inspection that would have shown how it performed. The 2021 external validation and the 2026 validation of the updated model are that check, arriving years late. The 2026 authors still urge local validation before trusting the model.

  • assumed

    This example does not model any patient or sepsis outcome. It shows how errors move among the model, the clinicians, and the record. The patients being scored are outside the network. The accuracy figures, STAT News's count of roughly 109 alerts per true case, the finding about undisclosed inputs, and the retune come from the case file. Nothing in this example computes them.

What this example does not show

Show all 2 limitations
  • This example does not show what happened to any patient, or whether anyone's sepsis was caught or missed. It shows how errors move among the model, the clinicians, and the record. The patients being scored are outside the network. The accuracy figures, STAT News's count of roughly 109 alerts per true case, and the retune come from the case file. Nothing in the network computes them.
  • The example opens already overloaded by choice, as the case documents. The model was switched on before anyone independent validated it, the alert volume was heavy, and Epic kept the model closed to outside inspection. The accuracy figures come from the external validation, and the alert ratio from STAT News's reporting. Neither is computed from the network.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and switched on across hundreds of hospitals was externally validated in 2021 across 38,455 hospitalizations at an academic health system: it achieved an area under the curve of 0.63, identified only 33 percent of sepsis cases, and had a positive predictive value of about 12 percent, generating roughly 109 alerts for every true sepsis case — a real-world performance the vendor had not fully examined before selling the model, and which an investigation attributed in part to undisclosed features such as antibiotic-order data that inflated internal validation.

    empirical
    • Academic Wong, A., Otles, E., et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307
    • Investigative STAT News (2021, July 26). Epic's AI algorithms, shielded from scrutiny by a corporate firewall, are delivering inaccurate information on seriously ill patients. https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/
  • After external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, and reducing its reliance on antibiotic-order features. A 2026 multicenter prospective validation of the updated model across 227,091 encounters reported an area under the curve of 0.82 to 0.92 with positive predictive value of 0.13 to 0.26 and substantial between-site variability, and its authors urged local validation and alert-silencing strategies rather than trusting the model out of the box — a correction that arrived only after independent scrutiny of a model that had already been deployed at scale behind a corporate firewall shielding it from outside inspection.

    empirical
    • Investigative STAT News (2022, Oct 3). Epic overhauls popular sepsis algorithm criticized for faulty alarms. https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/
    • Peer-reviewed Wong, A., Currey, D., Schwinne, M., et al. (2026). Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model. JAMA Network Open. https://doi.org/10.1001/jamanetworkopen.2026.0181 https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2845595

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical decision support & deterioration alerting domain page.

Levers available here and the patterns behind them

Documented case histories