Skip to content

PAN Lab example

TREWS sepsis early-warning system

The alert that works only when confirmed: a sepsis early-warning model

TREWS, built at Johns Hopkins, scores inpatients at five Hopkins hospitals for sepsis. Its benefit came only from alerts a provider confirmed within three hours.

See more

The Targeted Real-Time Early Warning System, TREWS, is a machine-learning model built at Johns Hopkins and run at five Johns Hopkins Medicine hospitals. It continuously scores hospitalized patients from their health records for sepsis, a life-threatening condition in which the body's response to illness damages its organs. When risk crosses a threshold, it alerts a provider to evaluate.

Where it was used and how it was studied

TREWS was deployed across the five hospitals of Johns Hopkins Medicine, in Maryland and Washington, DC. A prospective study followed 590,736 monitored patients as their care happened. It was published in Nature Medicine in 2022.

The case file calls it the largest prospective evaluation of a machine-learning sepsis system on record.

What the study found

The finding depends on a condition, and the condition is the provider's confirmation. Among patients who went on to have sepsis, some had their alert evaluated and confirmed by a provider within three hours. Their in-hospital mortality was 3.3 percentage points lower, a relative reduction of 18.7 percent, after statistical adjustment. The case file does not name the comparison group.

They also had less organ failure and shorter hospital stays. The benefit did not come from the alert firing. It came from the alert being confirmed and acted on in time.

Why confirmation varied

A companion paper on adoption, also published in Nature Medicine in 2022, looked at why providers did or did not confirm alerts. It found that confirmation varied with the provider's experience of the system, the unit's culture, and the context in which the alert arrived.

So the step that carried the whole benefit was itself uneven. The case file reads it as driven by trust and workflow rather than automatic.

What the evidence can and cannot claim

The evaluation is observational, not randomized. The mortality benefit is an association between confirmation and outcome. Providers who engaged with alerts may differ from those who did not, in caseload, in the patients they saw, and in attentiveness. The statistical adjustment cannot fully capture that. Confirmation may partly mark patients who were already going to do better.

The evaluation is also developer-led. TREWS was built at the deploying institution and commercialized through a company founded by its principal investigator. That company is Bayesian Health. The strongest numbers were produced on the developer's own patients, by the party with the greatest stake in the result.

The study is prospective and peer-reviewed, which is stronger evidence than most deployed clinical AI can show. It is not independent. At the time of the evaluation, no outside group had replicated the mortality effect. Peer review of the two papers is the only outside check in the record.

What an accuracy figure does not show

Two things matter here that a model's accuracy figure cannot show. The first is whether the confirmation step stays staffed and working. The benefit could be lost without touching the model, if confirmation wears away under caseload. It could also be lost at a site where trust and unit culture lead providers to stop looking.

The second is who validated the model. The strongest evidence here was produced by the people who built and commercialized it.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A pathway counts as closed once it passes on only a few mistakes. It need not stop them all.

This case's budget is 11 units. It is the cost of three changes the record supports. They are keeping provider confirmation staffed, reviewing the deployment on a fixed rhythm, and adding an independent check on TREWS. Each tool costs units from this budget.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met within the budget.

Under Service and Safety Targets and All Governance Targets, the available tools cannot fully address this deployment. Both levels require you to close every failure pathway. One pathway stays open even with all eight tools at their highest settings: a provider confirming or dismissing the alert. None of the tools offered here acts on it.

That pathway is the step that carried the whole measured benefit. Closing it would mean no provider evaluates the alerts. This is a finding about the deployment, not a gap in your approach.

Stylized model of a documented deploymentClinical decision support & deterioration alerting

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the TREWS-class sepsis early-warning alert network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 7 assumptions
  • assumed

    This network and the Lab's Proprietary Sepsis Model example share one structure. Both records describe the same kind of system. A model continuously scores live patient data inside the electronic health record and alerts the treating provider. Both records also include a role for outside validation of the model. Neither record documents a part the other lacks, so drawing a difference would be invention. The live data stream is drawn as its own input because it is this record's defining input. The system monitored 590,736 patients, scoring vital signs and lab results as they were recorded. The two examples differ in how they behave. Here, providers confirm alerts within three hours, and a prospective evaluation was carried out and published. The other example faces a flood of alerts and a dispute over how much its vendor discloses about the model.

  • assumed

    The benefit here depends on one pathway. The prospective multi-site study found lower mortality among sepsis patients whose alert a provider confirmed within three hours. So the provider's confirmation carries the benefit, not the model. This network treats that confirmation as real but uneven. The companion adoption paper documents that it varies with provider experience, unit culture, and alert context. This network assumes providers' workload and staffing are in balance, because this deployment is not documented as overwhelmed. What varies here is engagement, not headcount. That is the honest difference between this example and the overwhelmed ones in its domain.

  • assumed

    This network counts the published evaluation as a real check on TREWS. A prospective, multi-site, peer-reviewed evaluation of 590,736 monitored patients exists, so treating the check as absent would misstate the record. This network does not count it as a full check either, because that evaluation is observational and developer-led. TREWS was built at the deploying institution and commercialized through a company founded by its principal investigator. A published self-evaluation is a real read of the system, but not an arm's-length one. The difference between the two is what an independent check would add.

  • assumed

    This network models the pattern documented in the TREWS case file: a sepsis alert whose benefit depends on confirmation. It is not a reconstruction of TREWS. Its defining feature is that the measured mortality benefit came entirely through providers confirming alerts within three hours. The alert on its own did not reduce mortality.

  • baseline

    Unlike most examples in this Lab, the human step here works: the provider's evaluate-and-confirm step carried the entire measured benefit. But a companion adoption study found that confirmation varied with provider experience, unit culture, and alert context. So the confirmation rate belongs to the workflow, not to the tool. That is why the benefit can be lost without any change to TREWS, if the confirmation step wears away.

  • baseline

    What this example lacks is independence. The published study of TREWS was prospective, multi-site, and peer-reviewed across 590,736 patients. But it was observational and developer-led, run by the party that built and commercialized the model. At the time of the evaluation, no outside group had replicated the mortality effect. Peer-reviewed and prospective is not the same as independently validated. The missing check is the one that developer-led evidence cannot supply about itself.

  • assumed

    No patient or clinical outcome is modeled here. The Lab follows how errors move between TREWS, providers, and the health record. The patients being scored are not part of it. The mortality benefit, its dependence on confirmation, and the developer-led and observational caveats come from the case file. None of them is computed from this network.

What this example does not show

Show all 2 limitations
  • No patient or clinical outcome is modeled. The Lab follows how errors move between TREWS, providers, and the health record, and the patients being scored are not part of it. The sepsis mortality benefit, its dependence on confirmation, and the developer-led and observational caveats come from the case file. None of them is computed on this network.
  • The measured benefit is an observational association between provider confirmation and outcome. No randomized trial, one that assigns patients by chance, has confirmed it as an effect of the algorithm. Providers who confirmed alerts may differ systematically from those who did not. So confirmation may partly mark patients who were already likely to do better.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospectively across five hospitals of an academic health system covering 590,736 monitored patients — the largest prospective study of an ML sepsis system on record. Its central finding was conditional on the human loop: sepsis patients whose alert was evaluated and confirmed by a provider within three hours had a 3.3 percentage-point absolute and 18.7 percent relative adjusted reduction in in-hospital mortality, with less organ failure and shorter stays, while the alert on its own did not; a companion study found provider uptake varied with experience, unit culture, and alert context.

    empirical
    • Peer-reviewed Adams, R., Henry, K.E., et al. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28(7), 1455-1460. https://doi.org/10.1038/s41591-022-01894-0 https://www.nature.com/articles/s41591-022-01894-0
    • Peer-reviewed Henry, K.E., et al. (2022). Factors driving provider adoption of the TREWS machine learning-based early warning system and its effects on sepsis treatment timing. Nature Medicine, 28. https://doi.org/10.1038/s41591-022-01895-z https://www.nature.com/articles/s41591-022-01895-z
  • The TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: it was built at the deploying institution and commercialized through a company founded by its principal investigator, and confirmation-associated benefit is an observational association rather than a randomized effect of the algorithm — providers who engaged with alerts may differ from those who did not in ways the adjustment does not capture. The strongest numbers in the record therefore come from the party with the strongest interest in them, and no independent replication of the mortality effect had been published.

    empirical
    • Peer-reviewed Adams, R., Henry, K.E., et al. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28(7), 1455-1460. https://doi.org/10.1038/s41591-022-01894-0 https://www.nature.com/articles/s41591-022-01894-0

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical decision support & deterioration alerting domain page.

Levers available here and the patterns behind them

Documented case histories