PAN Lab example
Kaiser Permanente Suicide-Risk Model
The added sensor: an EHR-embedded suicide-risk score
Kaiser Permanente's machine-learning score flags intake patients at risk of a suicide attempt. The flag, or a positive questionnaire, prompts a clinician to assess them.
See more
Kaiser Permanente Northern California's suicide-risk model is a machine-learning model built into its electronic health record. It scores patients at intake to a large virtual mental-health program. When a score passes a pre-set threshold, a dashboard flag reminds the intake clinician to assess suicide risk.
How it is used
The program handles more than 5,000 intake visits a month. An internal predictive-analytics core computes the score in near real time. Only set kinds of visits trigger a score, and it is ready about 30 minutes later.
The flag joins an alert that already existed: a positive self-report screen. The screen means two questionnaires, the PHQ-9 (Patient Health Questionnaire-9) and the C-SSRS (Columbia-Suicide Severity Rating Scale).
Either alert starts the same workflow. The clinician is reminded to conduct a suicide-risk assessment at the upcoming visit. If the patient does not attend, up to two more attempts are made to contact them (outreach).
The team chose intake as the point of care because, by its own account, over 70 percent of treatment disengagement happens after the first or second visit. The sources read for this case do not corroborate that figure independently.
Who decides
The flag decides nothing. It is advisory and never gates care. The clinician conducts the suicide-risk assessment and decides any response.
The sources publish no rate at which clinicians accept or override the flag, or complete the assessment.
What the added flag changes
Either alert alone starts the workflow. So the model can add patients to the workflow but can never remove one. The two alerts are added together, not compared.
The event the model predicts is rare. In the validation study, 0.17 percent of intake appointments were followed by a suicide attempt within 90 days. Among the tenth of appointments the model ranked riskiest, 0.8 percent were.
So the large majority of flagged patients will not attempt in that window. At more than 5,000 intakes a month, the flag adds many false alarms and much clinician work. The implementation team itself raised that caution.
What the validation found
Papini and colleagues published the validation in JAMA Psychiatry in 2024. They studied 1,623,232 intake appointments from 835,616 patients, from January 2012 to April 2022. Of those appointments, 2,800, or 0.17 percent, were followed by a suicide attempt within 1 to 90 days. Of those attempts, 78, or 3.9 percent, were fatal.
The model's area under the curve was 0.77. That measure shows how well a score ranks a patient who later attempts above one who does not, where 0.5 is chance and 1 is perfect.
Sensitivity is the share of real cases the model flags. Specificity is the share of other patients it leaves unflagged. Sensitivity was 37.2 percent at 95 percent specificity, and 18.8 percent at 99 percent specificity.
The tenth of appointments ranked riskiest held 48.8 percent of the appointments later followed by an attempt. Its positive predictive value, the share of those appointments followed by an attempt, was 0.8 percent.
Ranking accuracy varied by group. The area under the curve ranged from 0.69 for one race group to 0.89 for another. It was 0.79 for women and 0.74 for men. The team cited this as an equity consideration.
TechTarget's HealthTech Analytics restated these figures in 2024. That was a restatement of the study, not a separate test.
How it was built
Kaiser Permanente's Division of Research and data-science staff at The Permanente Medical Group built and validated the model. It draws on earlier work by the Mental Health Research Network.
The implementation followed a five-step community-engaged process. First, the team assessed whether the model worked as well for each demographic group. Second, it interviewed patients with lived experience of suicide-related thoughts. Third, it gathered clinicians' views. Fourth, it assembled the workflow with operational leaders. Fifth, it educated and tested, in repeated rounds.
An institutional review board, which is an ethics review board, approved the work. It ran under a Delivery Science Grant. The authors stress that deployment must respect patient preferences, avoid reinforcing inequities, and manage workforce demands.
What the evidence does and does not show
The published record is a feasibility and design report, not an effectiveness study. It documents the workflow and its rationale. It presents no evaluation showing that the deployment reduced suicide attempts.
The peer-reviewed case study appeared in NEJM Catalyst Innovations in Care Delivery in February 2026. A medRxiv preprint from April 2025, which is not peer reviewed, mirrors it.
A related Mental Health Research Network study across four health systems found that models of this kind kept their ranking accuracy despite changes in how care was delivered.
Which record system
Kaiser Permanente Northern California runs KP HealthConnect, an Epic system, and the sources describe dashboards built within the record. The peer-reviewed text says electronic health record and names no vendor. So naming Epic is well-supported context, not a claim the study makes.
No large language model or generative AI is involved. This is a regression model over structured record fields. Reading it as a generative AI tool overstates it.
Where to look
Watch the two alerts that meet at the intake clinician: Model flag to the clinician's dashboard and Screen alert to the clinician. The case file's defining gap is that no standing check compares them.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway does not stop the work along it. It means mistakes stop passing along it.
Self-correcting means the network's mistakes are corrected rather than building on each other. At a tipping point, they are close to building on each other.
This case has a budget of 9 units. Explore (No Targets) sets no targets.
Under Service Targets Only, one tool is enough to meet the targets. The cheapest each cost 2 of the 9 units: Keep skills sharp, Escalate checks, Review on schedule, or Mark AI-written records.
Under Service and Safety Targets and All Governance Targets, the targets can also be met. Both levels ask you to close every failure pathway, among other targets. The cheapest way costs 7 of the 9 units and uses three tools: Escalate checks, Mark AI-written records, and Store less data.
At both levels, eight distinct combinations fit the budget, counting stronger settings. Every one of them includes those three tools.
Escalate checks stops mistakes passing along Model flag to the clinician's dashboard. Mark AI-written records stops them passing along four pathways. They are History read at the assessment, Record data read by the model, Screen alert to the clinician, and Screen answers as model input. Store less data stops them passing along Assessments written to the health record.
More is not better here. Every tool at its standard setting costs 21 units, and at its strongest setting 36. Either way every failure pathway closes, but the model is no longer clearly helping the work. So using everything misses the targets at every level that sets them.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Kaiser-EHR-class embedded suicide-risk flag network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 6 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
This network follows the pattern documented in the Kaiser Permanente suicide-risk case file: a risk flag built into the electronic health record and its workflow. It does not rebuild the actual model, dashboards, or workflow.
- baseline
The network's defining feature is how two alerts combine. The model's flag was added to an existing self-report screen: the PHQ-9 (Patient Health Questionnaire-9) and the C-SSRS (Columbia-Suicide Severity Rating Scale). Either alert starts the same assessment-and-outreach workflow. So the added flag can raise the number of flagged patients but never lower it. The network draws both alerts reaching one clinician workflow, combined rather than compared. It includes the check that would compare them, which the sources do not describe in use.
- assumed
The network assumes pathways between peers of both kinds. Clinicians pick up from each other assessment habits and how much weight to give the flag. One score repeats its blind spots across every intake: its ranking accuracy ranged from 0.69 to 0.89 across race groups. Against these, a real, staffed check ran before launch. The design team reviewed whether the model worked as well for each demographic group, and clinician-managers tested the workflow in repeated rounds.
- assumed
The network's defining gap is that the two alerts are combined, not compared, and no standing monitoring runs after launch. It shows this as a check on the model that the sources do not describe in use. The pre-launch subgroup review was real. It was not carried into continuous monitoring of who the flag reaches, who acts on it, or how the false alarms are absorbed. No rate of clinician flag acceptance, override, assessment completion, or workload is public.
- assumed
The flag is advisory and never gates care. The clinician keeps discretion over the assessment and any response throughout, so the network treats the clinician's correction as genuine. No public override or acceptance rate exists, so how much the clinician corrects is an assumption, not a measured figure.
- assumed
The Combined risk flag marks where the two alerts join before they reach the clinician. No pathway starts or ends at it, and it does not change how mistakes pass along the network. It shows that the two alerts are added together, not compared.
- assumed
The measured differences in the model's ranking accuracy across demographic subgroups are recorded in the case file as outside observations. This network follows how mistakes pass between the model, the staff, and the records, not demographics. It estimates no differential harm, and no suicide or crisis outcome for the patients served.
What this example does not show
Show all 3 limitations
- This example never models suicide or any crisis outcome. It follows how mistakes pass between the model, the staff, and the records. The patients the flag serves are not in the network. Here a flag, an assessment, or an outreach attempt is a signal inside an institution, never a life. The reports on this deployment focus on feasibility and design. They present no evidence that it reduced suicide attempts. The case file records that boundary, and any real outcome would be measured outside a network like this one.
- The case file records the measured differences in the model's ranking accuracy across subgroups as outside observations. The area under the curve ranged from 0.69 to 0.89 across race groups, and was 0.79 for women and 0.74 for men. Estimates for small subgroups were uncertain. This example asserts no flag rate for any subgroup, and estimates no differential harm to the people served.
- Naming the record as KP HealthConnect, an Epic system, is well-supported context: Kaiser Permanente Northern California runs it. The peer-reviewed record says electronic health record and names no vendor. No large language model or generative AI is involved. This is a regression model over structured record data, and it should not be read as a generative AI tool.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Kaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electronic health record of a large virtual mental-health program that handles more than 5,000 intake visits a month; the model is scored in near-real-time (about a 30-minute delay after an encounter trigger) and, at pre-set thresholds, flags high-risk patients to the intake clinician, routing them into the same suicide-risk-assessment and outreach workflow that a positive self-report screen (the PHQ-9 and Columbia-Suicide Severity Rating Scale) triggers, so the machine flag and the self-report alert are effectively OR-merged. In a study of 1,623,232 intake appointments (2012 to 2022, base rate 0.17 percent) the model reached an area under the ROC curve of 0.77 and its top risk decile captured 48.8 percent of appointments later followed by an attempt, but with a positive predictive value of about 0.8 percent.
empirical- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298
- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full
- Academic Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/
Because the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake is very low (0.17 percent) and the positive predictive value in the top risk decile is about 0.8 percent, the large majority of flagged patients will not attempt suicide in the window, so adding the machine-learning flag as a redundant sensor OR-merged onto the existing self-report screen imports a substantial false-positive and clinician-workload burden at scale — a caution the implementation team itself raised. The implementation reports are feasibility- and design-focused and present no evaluation showing the deployment reduced suicide attempts.
empirical- Academic Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/
- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298
- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Keep skills sharp — Deskilling-arrest mandate
- Escalate checks — State-feedback vigilance
- Check with a second model — Cross-model verification
- Review on schedule — Oversight cadence & retrospectives
- Require sign-off — Conformity assessment gate
- Mark AI-written records — Provenance labeling
- Upgrade model — Improve the model
- Store less data — Data minimization
Documented case histories
- Kaiser Permanente Suicide-Risk Model
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Crisis Text Line & Loris.ai
- LyssnCrisis counselor QA at ProtoCall Services (988)
- NarxCare
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- Tessa chatbot replacing the NEDA eating-disorder helpline
- Woebot (a governed app wind-down)