PAN Lab example
Advance Alert Monitor (AAM) deterioration model
The alert that never reaches the bedside: a screened deterioration model
Kaiser Permanente's Advance Alert Monitor flags hospital patients predicted to worsen in about twelve hours. Regional nurses screen every alert before the bedside team acts.
See more
The Advance Alert Monitor is a predictive model that Kaiser Permanente Northern California built and runs inside its own electronic health record. Every hour, around the clock, it scores adult inpatients at 21 hospitals. It raises an alert roughly twelve hours before it predicts a patient's condition will worsen.
Where the alert goes
The alert does not go to the bedside. It goes to a regional team of critical-care nurses, called virtual quality nurse consultants. The sources call the team virtual, and do not say where its nurses work.
They screen every alert around the clock and work up the patient's chart, reviewing it in detail. Only then do they escalate an alert to the on-site rapid-response team. A rapid-response team is a hospital team that comes to a patient's bedside when their condition worsens.
So the design has two tiers: the regional screening team first, then the bedside team. A 2022 paper in the Joint Commission Journal on Quality and Patient Safety describes this way of operating.
What the screening team does
The screening team does two jobs at once. It is an oversight step: it screens out the model's false alarms, so the bedside does not see them. It is also a standing cost: a regional nursing service, staffed around the clock, that exists only to handle the model's alerts.
That the health system built such a team says something about the raw alerts. A system that screens every alert before anyone acts is saying the unscreened alerts hold too many false alarms to act on directly.
What the evaluation found
A 2020 evaluation in the New England Journal of Medicine, by Escobar and colleagues, studied the program. It associated the alert-driven rapid-response workflow with lower mortality across the deployment. The finding is an observational association, not the result of a randomized trial.
The finding belongs to the model and the screening team together, not to the model alone. Without the screening team, nothing remains to produce the measured effect. Quoting the benefit while picturing the model alone quotes a number nobody measured.
The share of alerts that proved correct was not published.
Who evaluated it
The model was built and evaluated inside the health system that deployed it. That gave the evaluators close access, and it limits their independence. The record holds no standing outside audit of the model, and no validation by a party with no stake in the result. A validation is a test of whether the model's alerts are accurate.
What is at stake
Two things matter here that the mortality figure cannot show. The first is that the benefit was measured with the whole staffed screening team in place.
The second is what happens if the team is cut to save money once the model is trusted. That would not make the tool cheaper to run safely. It would remove what the benefit ran through and push the false alarms onto the bedside.
The team's own staffing is the real limit. With too few nurses for the number of alerts, screening can become a rubber stamp. The two tiers would then turn back into a flood of unscreened alerts.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest way to meet them is a single tool costing 2 of this case's 11 budget units.
Under Service and Safety Targets and under All Governance Targets, this case is not fully addressable with the available tools. Both levels ask you to close every failure pathway, among other targets. Using all eight tools at their standard settings costs 18, well over the budget, and two failure pathways stay open. At their strongest settings the tools cost 28, and the same two stay open.
They are the nurse consultants working up each alert against the chart, and the nurse consultants escalating a screened alert to the bedside team. These two pathways are the screening team's own work. Every tool offered here, used together at full strength, leaves both open. The benefit the evaluation measured ran through that screening team. This is a finding about the deployment, not a gap in your approach.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the AAM-class deterioration alert with a screening tier network: 6 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
The regional alert queue is shown as a work list. Alerts from all 21 hospitals go to one regional team of virtual quality nurse consultants, as the 2022 program description sets out. That team screens every alert around the clock, works up the chart, and only then escalates to the on-site team. The queue is what lets one team serve a whole region. Without it, the two tiers would read as two separate desks, not one pooled screening team.
- assumed
This network assumes the deployment has more screening staff than its workload needs. That is unusual, and it is the opposite of how the Lab shows Duke Health's Sepsis Watch, another hospital alert case. The published evaluation ties the benefit to the whole two-tier staffing design, not to the model. So a dedicated round-the-clock screening team is the standing cost the benefit depends on. The network also gives the model some credit for outside checking, because its evaluation passed peer review. That peer-reviewed evaluation appeared in the New England Journal of Medicine. Such an evaluation is rare among the Lab's cases, and it sets this case apart from those resting on a vendor's own reports.
- baseline
This network follows the two-tier screening design documented in the Advance Alert Monitor case file. It does not rebuild the actual model. Its defining feature is where the alert goes: to a dedicated regional team of virtual nurse consultants, not to the bedside. The lower mortality the evaluation found belongs to the model and the screening team together, not the model alone.
- assumed
The screening team is assumed to be active both in receiving alerts and in escalating them. It is a staffed, round-the-clock job. It carried the benefit by screening out the model's false alarms before the bedside saw them. Its capacity is the real limit. With too few nurses for the number of alerts, the two tiers would turn back into a flood of unscreened alerts. So the benefit can be lost by cutting the team, without changing the model.
- assumed
This network assumes no validation by a party outside the health system exists. The evaluation of record was run inside the deploying health system, so the strongest evidence comes from the developer's side. The record holds no validation by a party with no stake in the result. The share of alerts that proved correct was not published. The dedicated screening team is itself the health system's answer to false alarms in the raw alerts.
- assumed
No patient or clinical outcome is modeled here. The Lab reads how mistakes pass between the health system's people, tools, and records, not what happens to patients. The patients being scored stay outside the network. The link to lower mortality and the standing staffing cost are described in the case file. Nothing in this diagram computes them.
What this example does not show
Show all 2 limitations
- No patient or clinical outcome is modeled. The Lab reads how mistakes pass between the health system's people, tools, and records. The patients being scored stay outside the network. The link to lower mortality and the standing staffing cost are described in the case file, and nothing on this diagram computes them.
- The lower mortality comes from an observational comparison, not a randomized trial. The evaluation was run inside the health system that built the model. The share of alerts that proved correct was not published. So the screening team's filtering of false alarms is a documented design choice, not a measured rate.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The Advance Alert Monitor is an in-hospital deterioration model running around the clock across 21 hospitals of an integrated health system, scoring inpatients hourly and firing roughly twelve hours before predicted deterioration; a 2020 New England Journal of Medicine evaluation associated its alert-driven rapid-response workflow with lower mortality. Its defining feature is where the alert goes: not to the bedside, but to a dedicated regional tier of critical-care virtual quality nurse consultants who screen every alert around the clock, work up the chart, and only then escalate to the on-site rapid-response team — so the measured benefit is priced against the whole two-tier staffing topology, not the model alone.
empirical- Academic Escobar, G.J., Liu, V.X., Schuler, A., Lawson, B., Greene, J.D., & Kipnis, P. (2020). Automated Identification of Adults at Risk for In-Hospital Clinical Deterioration. New England Journal of Medicine, 383(20), 1951-1960. https://doi.org/10.1056/NEJMsa2001090 https://www.nejm.org/doi/full/10.1056/NEJMsa2001090
- Academic The Kaiser Permanente Northern California Advance Alert Monitor Program: An Automated Early Warning System for Adults at Risk for In-Hospital Clinical Deterioration (2022). Joint Commission Journal on Quality and Patient Safety. https://www.jointcommissionjournal.com/article/S1553-7250(22)00110-6/fulltext
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
All of them in context on the Clinical decision support & deterioration alerting domain page.
Levers available here and the patterns behind them
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Upgrade model — Improve the model
- Check with a second model — Cross-model verification
- Mark AI-written records — Provenance labeling
- Store less data — Data minimization
Documented case histories
- Advance Alert Monitor (AAM) deterioration model
- TREWS sepsis early-warning system
- Sepsis Watch deep-learning detection system
- Proprietary EHR sepsis model (external validation)
- nH Predict Utilization Review
- Cost-Proxy Care Stratification
- CA-CDS Child Abuse Alerting
- IDx-DR Autonomous Screening
- Viz.ai LVO Stroke Triage
- IBM Watson for Oncology
- OPTN eGFR Waiting-Time Correction
- Practice Fusion Pain CDS
- UBH Level of Care Guidelines (Wit v. UBH)
- EviCore by Evernorth: the review threshold
- Cigna PxDx