Skip to content

PAN Lab example

REACH VET

The flag that works and the outcome it misses: a suicide-risk model

Each month REACH VET, a Veterans Affairs model, flags high-risk veterans for suicide-prevention outreach. Appointments rose, but two evaluations found no drop in suicide deaths.

See more

REACH VET is a suicide-prevention program the Veterans Health Administration (VHA) has run nationwide since April 2017. Each month its statistical model scores every patient who used VHA care in the prior two years, using their health records. It flags the 0.1% at highest estimated risk at each facility, so a clinician can re-evaluate them and reach out.

How the model works

REACH VET stands for Recovery Engagement and Coordination for Health, Veterans Enhanced Treatment. VHA is the health-care arm of the Department of Veterans Affairs (VA).

The model is a penalized regression, known as Lasso: a statistical method that shrinks the weight of variables adding little to a prediction. Variables are the pieces of patient information it weighs. It uses 61 of them in six groups. The first four are demographics (including age, gender, marital status, and race or ethnicity), diagnoses, medications, and use of health services. The last two are prior suicide attempt and combinations such as marital status by gender. It was cut down from a 2015 proof-of-concept with 381 variables, which was too unwieldy to compute. Ronald Kessler of Harvard refined it and added machine-learning methods.

Each month it scores about 6.28 million people and lists the top 0.1% at each VA facility. That is roughly 6,300 to 6,700 veterans a month, and more than 130,000 since 2017.

Who decides

A facility Suicide Prevention Coordinator receives each flagged name on an internal dashboard and notifies the veteran's clinician. The clinician re-evaluates the veteran's suicide risk and treatment. They then reach out without a script, offering enhanced care, safety planning, closer monitoring, and support with coping.

The flag is advisory. VA frames the model as a prompt for clinician review, not an automated decision, and states that AI will "never replace human intervention." Veterans may decline the outreach.

So this human review is real and staffed, in two stages: the coordinator, then the clinician.

What the flag catches and misses

The flagged group is at far higher risk. It dies by suicide at roughly 19 to 30 times the overall VHA rate.

But suicide is rare: about 179 a month among 6.28 million scored patients. An independent re-analysis of 2018 data found that about 0.05% of flagged veterans died by suicide. The flag missed about 98% of suicides, catching about 2%. For suicide attempts and deaths combined, about 5% of flagged veterans had the outcome.

Predictive value here means the share of flagged veterans who had the outcome. A separate retraining experiment, for veterans involved in the legal system, changed the predictive value for the combined outcome only slightly, and not significantly. That suggests the limit is the rarity of suicide itself, not a modeling flaw that could be fixed.

What the evaluations found

VA's own 2021 evaluation compared 173,313 veterans across 141 facilities. It used triple differences, a statistical method that compares changes over time across several groups. That is an observational design, not a randomized trial. REACH VET was associated with more completed outpatient appointments, more new safety plans, fewer mental-health admissions, and fewer documented suicide attempts. It was not associated with fewer deaths by suicide, or fewer deaths from any cause.

A 2025 follow-up of 266,246 observations found the same: no reliable difference in deaths. The program improved what it could measure quickly. The outcome it was built to change did not change in either study.

The 2024 investigation

In May 2024 The Markup and The Fuller Project, two news organizations, published an investigation. It reported that the model treated being a white man as a stronger risk signal than factors affecting women. It reported that the model gave weight to veterans who were "divorced and male" or "widowed and male," but to no female group.

The investigation reported that military sexual trauma and intimate-partner violence were excluded from the model. Both are linked to higher suicide risk in women veterans. It also reported that the model did not account for LGBTQ+ identity. VA research finds transgender veterans die by suicide at about twice the rate of veterans who are not transgender.

VA's suicide-prevention executive framed the exclusions as a judgment about predictive strength. The executive said military sexual trauma "was not among the most powerful for us to be able to predict suicide risk." So the finding is documented but contested, not a design intent VA has confirmed.

The female-veteran suicide rate rose about 24% from 2020 to 2021. That is roughly four times the rise among male veterans.

REACH VET 2.0

In October 2024 VA announced REACH VET 2.0. It adds military sexual trauma, intimate-partner violence, and medical factors specific to women. VA committed to evaluate it "for performance and bias before it is deployed." The announcement followed a bill from Senator Jon Tester requiring those changes.

By December 2025, trade press reported that 2.0 had launched in 2025 with the new factors. It also reported that race and ethnicity had been removed as variables. No peer-reviewed description of 2.0 or independent bias audit had been published, so its performance and fairness are not independently verified.

Oversight and support

The 2019 Hannon Act, a federal law, required a review by the Government Accountability Office (GAO). GAO's 2022 review described the program. It noted that VHA planned, but had not yet completed, studies of how the model's performance varied by age, sex, and race. It made no recommendations specific to REACH VET.

Congress has kept backing the approach. The appropriations law for military construction and VA for fiscal year 2026, signed in November 2025, funds suicide prevention and encourages expanded predictive modeling.

What this network is drawn from

This network follows the pattern the case file describes. It is not a reconstruction of the actual tool. It shows the model, the monthly flag list, the coordinator, the clinician, the health record, the mortality repository, and the program's evaluators. The veterans served, and whether they live or die, are outside the network.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Here a mistake is, for example, a veteran at real risk ranked off the list. Closing a failure pathway means mistakes stop passing along it. The work along it goes on: the coordinator still notifies the clinician.

This case has a budget of 11 units. Each tool costs the same at every target level.

Before any tool is used, mistakes are copied about as fast as they are corrected. Six failure pathways are open. They are Monthly flag list to coordinator, Coordinator notifies clinician, and Clinician weighs the flag. They are also Care actions written to the health record, Health record supplies the model's variables, and Clinician reads the veteran's history.

Three checks do nothing at the start: Pre-deployment subgroup and bias check, Care measures checked against deaths, and Program review by subgroup.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets are not met before any tool is used. That level asks for mistakes to be caught faster than they are copied, and for REACH VET to be helping the work.

Mark AI-written records meets that level's targets on its own, for 2 units. It closes Health record supplies the model's variables and Clinician reads the veteran's history. Peer sharing rules does too at its stronger setting, for 3 units. Lingering effects is a Dynamics setting, on by default, in which damage outlasts its cause. With it off, the standard setting of Peer sharing rules also does, for 2 units. In all, more than 3,600 sets of tools within the budget meet them, counting each stronger setting as a set of its own.

Under Service and Safety Targets and All Governance Targets, the targets can be met, but only by spending the whole budget. Both levels ask you to close every failure pathway, among other targets. Closing all six takes five tools.

Escalate checks closes Monthly flag list to coordinator. Peer sharing rules, which sets rules for what people pass to each other, closes Coordinator notifies clinician. It also starts the check named Program review by subgroup.

Keep prompts neutral closes Clinician weighs the flag. In the Lab it keeps questions put to an AI neutral. REACH VET takes no questions, so here it stands for keeping the clinician's own re-evaluation independent of the flag.

Mark AI-written records closes the two pathways from the health record, as above. In the Lab it marks machine-written content in the records so readers can weigh it. Here that means marking the model's risk flag wherever it is logged in the health record.

Gate record entries, which requires sign-off before anything enters the records, closes Care actions written to the health record. Store less data closes it instead.

So exactly two sets of tools meet the targets at these levels, and each costs all 11 units. They differ only in Gate record entries or Store less data.

Neither set includes the tools that start the two other checks. Check with a second model starts Pre-deployment subgroup and bias check, but the work then falls behind demand. Check copied records starts Care measures checked against deaths, but closes no failure pathway.

Using every tool on offer, each at its strongest setting, costs 46 units, more than four times the budget. It closes every failure pathway and starts all three checks. But the benefit REACH VET adds would fall too far, so the targets would be met at no level.

Stylized model of a documented deploymentBehavioral-health & crisis triage

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the REACH VET-class national suicide-risk flag network: 7 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 7 assumptions
  • assumed

    This network follows the pattern the REACH VET case file describes: a national suicide-risk flag that coordinators pass to clinicians. It is not a reconstruction of the actual tool. The deployment is live. The model has run monthly across the Veterans Health Administration (VHA) since April 2017. A recalibrated version launched in 2025, and Congress has encouraged expanding such predictive modeling.

  • baseline

    The network hinges on two checks that do nothing at the start: Pre-deployment subgroup and bias check, and Program review by subgroup. No independent check of how the model performed for different groups ran before it scored the whole VHA population. The Government Accountability Office's 2022 review found VHA had planned such studies but not completed them. A 2024 investigation reported that women were under-flagged, because military sexual trauma, intimate-partner violence, and LGBTQ+ identity were left out of the model. VA contests that reading. The gap went uncaught until that outside investigation and a recalibration years later, REACH VET 2.0.

  • assumed

    Unlike most networks in the Lab, this one draws the human review as working. The coordinator's notice to the clinician and the clinician's own judgment form a real, staffed, two-stage review. The flag is advisory and never sets off an automated care action, and veterans may decline outreach. No public figure exists for how often clinicians disagree with the flag or veterans decline. So the network's picture of that review rests on a judgment, not a measurement.

  • baseline

    The network draws the gap between care measures and suicide deaths as a comparison between two record stores that does nothing at the start. The electronic health record holds the near-term measures the program improved: completed appointments, new safety plans, and fewer documented attempts. A separate mortality repository holds the outcome the outreach did not move, death by suicide. Two Department of Veterans Affairs (VA) evaluations, in 2021 and 2025, found the improvement without a reduction in suicide deaths. In the sources, that comparison happens only in periodic studies, not routinely.

  • baseline

    The network places the demographic miss on the pathway where the health record supplies the model's variables, which is where the sources place it. VA gives enhanced outreach to a fixed 0.1% of patients, ranked by predicted risk. So the choice of variables decides who is eligible, and factors specific to women were left out of the original set. That is why a more accurate model would not have fixed it. A retraining experiment, for veterans involved in the legal system, barely changed the model's predictive value, the share of flagged veterans who had the outcome. The change that mattered was the variable set, and so who is eligible, not the accuracy.

  • assumed

    The network draws the monthly flag list as a step between the model and the coordinator. What defines it is its fixed size, 0.1% of patients at each facility, which the network treats as keeping the list short enough to act on. It has no effect of its own on how mistakes pass between parts of the network.

  • assumed

    No suicide or crisis outcome is part of this network. It follows only how mistakes pass between staff, tools, and records, and the veterans served are not in it. The case file records a sex disparity among the veterans served. From 2020 to 2021 the female-veteran suicide rate rose about four times as much as the male rate. The investigation's finding about the model is contested, because VA framed the excluded factors as simply less predictive. Nothing in this network computes it.

What this example does not show

Show all 3 limitations
  • This network never models suicide or any crisis outcome, and the veterans this flag serves are not in it. A flag or an act of outreach here stands for work between staff and systems, never a life. Two VA evaluations found more kept appointments and no reduction in suicide deaths. Those findings are recorded in the case file and were measured outside this network.
  • The demographic miss comes from a 2024 investigation by The Markup and The Fuller Project. It reported that the model treated being a white man as a stronger risk signal than factors specific to women. It also reported that military sexual trauma, intimate-partner violence, and LGBTQ+ identity were left out. VA contests this, framing the excluded factors as simply less predictive. The case file carries the finding as documented but contested. This network states no flag rate for any group, because the public figures by group are partial and contested.
  • An independent re-analysis of 2018 data found the flag caught about 2% of suicide deaths and missed about 98%. Those figures are for death by suicide. Do not confuse them with the higher predictive values, the share of flagged veterans who had the outcome, reported for suicide attempts and deaths combined. REACH VET 2.0, launched in 2025, had not been independently evaluated for performance or bias when the sources were gathered. So this network shows the documented structure of the first version, not a claim about 2.0.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Veterans Health Administration since 2017, scoring about 6.28 million patients and flagging the top 0.1% at each facility (roughly 6,300 to 6,700 veterans a month, more than 130,000 since 2017); an independent re-analysis of 2018 data found the top-0.1% flag has a positive predictive value near 0.05% and a false-negative rate of about 98% for death by suicide, and a 2024 investigation reported that the model treated being a white man as a stronger risk signal than factors specific to women and excluded military sexual trauma and intimate-partner violence from its variables, a characterization VA has contested by framing the excluded factors as less predictive.

    empirical
    • Academic Harris, Finlay, Meerwijk, Evaluating the accuracy of the VHA REACH VET suicide prediction model for legal involved veterans (npj Mental Health Research, 2025;4:53) https://pmc.ncbi.nlm.nih.gov/articles/PMC12535588/
    • Investigative Glantz, V.A. Uses a Suicide Prevention Algorithm to Decide Who Gets Extra Help. It Favors White Men. (The Markup with The Fuller Project, 2024) https://themarkup.org/news/2024/05/30/v-a-uses-a-suicide-prevention-algorithm-to-decide-who-gets-extra-help-it-favors-white-men
    • Trade press Graham, Inside VA's yearslong AI effort to uncover veterans at high risk of suicide (Nextgov/FCW, 2025) https://www.nextgov.com/artificial-intelligence/2025/07/inside-vas-yearslong-ai-effort-uncover-veterans-high-risk-suicide/406781/
    • Government U.S. Government Accountability Office, Veteran Suicide: VA Efforts to Identify Veterans at Risk through Analysis of Health Record Information (GAO-22-105165, 2022) https://www.gao.gov/assets/gao-22-105165.pdf
  • Two Veterans Health Administration evaluations of REACH VET found the program associated with improved proximal outcomes — more completed outpatient appointments, more new safety plans, and fewer documented suicide attempts — but not with reduced death by suicide: a 2021 triple-differences study of 173,313 veterans across 141 facilities found no association with suicide or all-cause mortality, and a 2025 follow-up of 266,246 observations replicated the null with all confidence intervals crossing one; both are observational rather than randomized studies.

    empirical
    • Academic McCarthy, Cooper, Dent et al., Evaluation of the REACH VET Suicide Risk Modeling Clinical Program in the Veterans Health Administration (JAMA Network Open, 2021;4(10):e2129900) https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2785078
    • Academic Dent, Cooper, McCarthy, The REACH VET Program and Mortality Outcomes Among Veterans at High Risk of Suicide (JAMA Network Open, 2025;8(7):e2519513) https://pmc.ncbi.nlm.nih.gov/articles/PMC12238888/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Behavioral-health & crisis triage domain page.

Levers available here and the patterns behind them

Documented case histories