Skip to content

PAN Lab example

Vanderbilt VSAIL suicide-risk alert

The alert that had to be dismissed: an EHR suicide-risk model

In a trial, Vanderbilt's suicide-risk alert led clinicians to screen in 42 percent of flagged visits as a pop-up, 4 percent as a chart icon.

See more

VSAIL, short for Vanderbilt Suicide Attempt and Ideation Likelihood, is a machine-learning model Vanderbilt University Medical Center built in house. It estimates a patient's risk of a suicide attempt within 30 days from routine electronic health records. In a trial, clinicians saw an alert at visits where that risk was 2 percent or higher.

What the trial tested

Vanderbilt ran a randomized trial, registered as NCT05312437 and deployed as Vanderbilt Safecourse. It ran from August 2022 to February 2023 in three Vanderbilt neurology clinics. Vanderbilt chose neurology because some neurological conditions carry a higher suicide risk.

VSAIL flagged 596 of 7,732 visits, about 8 percent, at a risk of 2 percent or higher. Each flagged visit was assigned by chance to one of two alert forms. One was a pop-up the clinician had to dismiss to go on. The other was a passive icon in the chart.

The score, the risk cutoff, and the suicide-risk screen offered, a standard set of questions, were the same in both. Only the form of the alert changed.

What it found

With the pop-up, clinicians chose to run a suicide-risk screen at 42 percent of flagged visits, 121 of 289. With the chart icon, they did so at 4 percent, 12 of 307.

The screen is the Columbia Suicide Severity Rating Scale, a standard set of questions a clinician asks the patient. Documented assessments on it rose to 22 percent, against 8 percent the year before.

Screening stayed the clinician's choice. The pop-up could be dismissed with no required action. About 58 percent of pop-up alerts and 96 percent of chart-icon alerts led to no screen.

What the trial could not show

No suicidal thoughts or attempts were documented in either group during 30 days of follow-up. The trial was explicitly not large enough to measure clinical outcomes. So it showed more screening, a process outcome, but could not speak to reduced harm. As the case file puts it, a gain in process is not a gain in outcome.

The tradeoff the researchers named

The researchers named alert fatigue as the central tradeoff. Alert fatigue is the reflex to dismiss a familiar interruption. Clinicians said they preferred the passive icon, which led to far fewer screens.

Most flags are false alarms. In a 2021 study, clinicians would need to screen 271 of the highest-risk patients to find one suicide attempt. They would need to screen 23 to find one patient with suicidal thoughts. This figure is called the number needed to screen.

Who decides

The clinician at the visit decides whether to run the screen. Vanderbilt's institutional review board, the committee that approves research on people, let the trial run without the usual individual consent. The sources do not say whose consent was waived. One reason given was to protect clinicians who might credibly disagree with VSAIL and decline to screen. The trial left them free to decline.

How the model was checked

The case file calls Vanderbilt's governance comparatively rigorous. Governance included review-board approval, embedded ethicists, and consultation with Vanderbilt's Office of Legal Affairs. The trial was registered in advance. A study of the model running silently, with no alerts, was published before any live alerting.

That silent study ran from June 2019 to April 2020, with 115,905 predictions on 77,973 patients. The model's early risk estimates were miscalibrated: the risks it predicted did not match the rates later seen. Vanderbilt recalibrated it by hand. The sources call this monitoring algorithmovigilance.

Area under the curve measures how well a score ranks a real case above a non-case, where 0.5 is a coin flip. In the silent study it was 0.797 for attempts across the medical center. In behavioral-health settings, where suicidality is most concentrated, it was 0.544.

VSAIL is not a device cleared by an independent regulator. A 2023 study applied it to 260,583 Navy primary-care patients at Naval Medical Center Portsmouth. Its area under the curve was 0.77 applied directly, and 0.92 after retraining.

Where to look

Watch the step from VSAIL to the clinician. The case offers a rare clean measurement: hold the score, the cutoff, the population, and the screen constant, and change only how the alert is delivered. Screening moved roughly tenfold.

What matters here is not the model's accuracy. It is how the alert is delivered, and the attention it commands.

The same step has a second edge. Force attention too bluntly with a flag that is mostly false alarms, and clinicians may learn to dismiss the very screen the alert exists to prompt.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A pathway counts as closed once only a few mistakes pass along it. The work along it goes on.

Contained means the network's mistakes are corrected rather than building on one another. Strained means the work the system supports is not keeping up.

This case has a budget of 10 units. It starts at the tipping point, where mistakes are on the edge of building on one another rather than being corrected. Three failure pathways start open: Record fields read by VSAIL, Screening recorded, and One model scores each patient.

Under Explore (No Targets), which sets no targets, almost any one tool keeps the mistakes contained. Under Service Targets Only, the same single tools also meet the targets. That level asks for the automated system to be helping the work.

The cheapest cost 2 units each: Mark AI-written records, Escalate checks, Keep skills sharp, Assign a challenger, and Peer sharing rules. Review on schedule does it alone only at its stronger setting, for 4 units.

Upgrade model does it alone at its standard setting, for 3 units, with lingering effects on, the setting you first see. With them off, it needs its stronger setting, for 5 units.

Understand the system costs 3 units under Explore and Service Targets Only, and 4 at the two higher levels. While it is on, four tools cost 1 unit less: Review the riskiest first, Assign a challenger, Peer sharing rules, and Upgrade model. At its stronger setting, Deep research, for 6 units, they cost 2 less, never below 1.

Under Service and Safety Targets and All Governance Targets, this case is not fully addressable with the available tools. Both levels ask you to close every failure pathway, keep the work from being strained, and have the system helping. All Governance Targets asks for it to be clearly helping.

Closing all three pathways fits the budget. The cheapest ways cost 7 of the 10 units: Mark AI-written records, Peer sharing rules, and either Store less data or Gate record entries. Every combination that closes all three includes those tools.

Within the budget, every such combination leaves the work strained, or leaves the automated system not helping the work. Adding a fourth tool that relieves the strain leaves the system not helping.

More money does not change that. With the budget set aside and every tool at its stronger setting, every failure pathway closes, and the automated system ends up hurting the work.

That is a finding about the deployment, not a gap in your approach.

Stylized model of a documented deploymentBehavioral-health & crisis triage

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the VSAIL-class EHR suicide-risk alert network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 6 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 8 assumptions
  • assumed

    This network follows the pattern the VSAIL case file documents: one suicide-risk alert, shown to clinicians in two forms, with very different results. It does not rebuild the actual model or its alerts.

  • baseline

    The network assumes the main step is the alert from VSAIL to the clinician. In the trial, making the same alert a pop-up instead of a chart icon raised screening from about 4 percent to about 42 percent of flagged visits. That is roughly tenfold. The score, the risk cutoff, and the screen stayed the same. The chart icon went unanswered far more often than not.

  • baseline

    Pushing the alert harder carries the risk the researchers named: alert fatigue. Most flags are false alarms. In the 2021 study, clinicians would need to screen 271 of the highest-risk patients to find one suicide attempt. A familiar pop-up on such flags invites dismissal by reflex, which wears down the screen the alert exists to prompt. The network treats the screen as the clinician's choice. Even the pop-up led to no screen at about 58 percent of flagged visits.

  • assumed

    The risk cutoff and the alert are drawn as parts of the step from VSAIL to the clinicians. Neither has links of its own in the network. Vanderbilt set the cutoff at a 2 percent risk, which flagged about 8 percent of visits. The case file reads that as a deliberate bound that keeps targeted screening survivable. Whether the bound is set right is a design question the case file records. The network does not compute it.

  • assumed

    The network assumes every score is computed from the health record, from records patients did not provide for this purpose. Race, a ZIP-based measure of neighborhood deprivation, and mental-health diagnoses are among the inputs. What those inputs mean for different groups of patients is covered in the last assumption, and stays outside the network.

  • assumed

    The monitoring part stands for Vanderbilt's own accuracy monitoring, which the sources call algorithmovigilance. The model ran silently before any alert. It was recalibrated by hand after its early predicted risks did not match the rates later seen. The network draws the monitoring as receiving VSAIL's scores. It does not check each alert, and it does not correct the model directly. Of the tools on offer, Review on schedule is the one that reviews VSAIL's controls on a fixed schedule. The case file calls this monitoring the only thing watching the model drift, a slow change in how accurate its scores are.

  • assumed

    The network assumes mistakes repeat in two ways. One model scores each patient it is run on, so its errors repeat, and it ranked risk worst in behavioral-health settings, where suicidality is most concentrated. Clinicians also pick up alert habits from one another. The network includes two checks the sources do not describe in use: a colleague's second read of a flagged visit, and an independent check of VSAIL's scores. So the slow monitoring is the only check the sources describe on a drifting model, one whose accuracy slowly changes.

  • assumed

    This network models no suicide or crisis outcome. It follows only how mistakes pass between VSAIL, the clinicians, and the record. The case file records how the model treats different groups of patients, measured outside a network like this one. The differences in the number needed to screen are reported but not settled as bias or as differences in underlying risk. No figure for harm to any group is given here.

What this example does not show

Show all 2 limitations
  • This example models no suicide or crisis outcome. It follows how mistakes pass between VSAIL, the clinicians, and the record. The trial measured a process outcome, whether a screen happened, not whether suicide was prevented. No suicidal thoughts or attempts were documented in either group during 30 days of follow-up. The trial was explicitly not large enough to measure clinical outcomes. So nothing here speaks to reduced harm. The case file records how the model treats different groups of patients, measured outside any network like this one.
  • The 42 percent and 4 percent figures come from one trial, in three Vanderbilt neurology clinics from August 2022 to February 2023. They do not come from routine alerting across the medical center. VSAIL is not a device cleared by an independent regulator. The number needed to screen and the 0.797, 0.836, and 0.544 accuracy figures come from a separate 2021 study, run with no alerts. Its reported differences between patient groups are not settled as bias or as differences in underlying risk.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In a single-center randomized trial across three Vanderbilt neurology clinics (August 2022 to February 2023), an EHR suicide-risk model flagged 596 of 7,732 encounters (about 8%) at a 2%-or-higher 30-day-risk threshold; making the identical alert interruptive rather than passive led clinicians to elect a suicide-risk screen in 42% of encounters (121/289) versus 4% (12/307) for a passive chart icon, an adjusted odds ratio of 17.70 (95% CI 6.42–48.79). Screening remained fully advisory: about 58% of interruptive and 96% of passive alerts produced no screening.

    empirical
    • Academic Walsh et al., Risk Model-Guided Clinical Decision Support for Suicide Screening: A Randomized Clinical Trial (JAMA Network Open, 2025; PMC11699529) https://pmc.ncbi.nlm.nih.gov/articles/PMC11699529/
    • Vendor AI tested for alerting clinicians of suicide risk at three VUMC clinics (Vanderbilt University Medical Center News, first-party institutional communication, 2025) https://news.vumc.org/2025/01/03/ai-tested-for-alerting-clinicians-of-suicide-risk-at-three-vumc-clinics/
    • Trade press Suicide prevention more feasible using AI-powered screening alerts (Healio Primary Care, 2025) https://www.healio.com/news/primary-care/20250122/suicide-prevention-more-feasible-using-aipowered-screening-alerts
  • In a separate 2021 prospective silent-mode study (115,905 predictions on 77,973 patients, June 2019 to April 2020), the model reported a c-statistic of 0.797 for suicide attempt and 0.836 for ideation center-wide but only 0.544 for attempt in behavioral-health settings, and in the highest-risk quantile the number-needed-to-screen was 271 for attempt and 23 for ideation. In the 2022 to 2023 trial no suicidal ideation or attempts were documented in either arm during 30-day follow-up, and the trial was explicitly not powered for clinical outcomes, so it measured a process outcome (screening) rather than reduced harm.

    empirical
    • Academic Walsh et al., Prospective Validation of an Electronic Health Record-Based, Real-Time Suicide Risk Model (JAMA Network Open, 2021; PMC7955273) https://pmc.ncbi.nlm.nih.gov/articles/PMC7955273/
    • Academic Walsh et al., Risk Model-Guided Clinical Decision Support for Suicide Screening: A Randomized Clinical Trial (JAMA Network Open, 2025; PMC11699529) https://pmc.ncbi.nlm.nih.gov/articles/PMC11699529/
    • Trade press Suicide prevention more feasible using AI-powered screening alerts (Healio Primary Care, 2025) https://www.healio.com/news/primary-care/20250122/suicide-prevention-more-feasible-using-aipowered-screening-alerts

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Behavioral-health & crisis triage domain page.

Levers available here and the patterns behind them

Documented case histories