Skip to content

PAN Lab example

REACH VET

The flag that works and the outcome it misses: a suicide-risk model

A national model flags the highest-risk 0.1% of patients each month; a coordinator routes each flag to a clinician, who re-evaluates and reaches out. Modeled on REACH VET. The human loop here actually works - so watch two things it cannot see: the vast majority of the target it misses by construction, and who its feature set leaves off the list.

Stylized model of a documented deploymentBehavioral-health & crisis triage

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the REACH VET-class national suicide-risk flag network: 7 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the national coordinator-mediated suicide-risk-flag pattern documented in the REACH VET case file - not a reconstruction of the actual tool. It is a live, expanding deployment: the model has run monthly across the Veterans Health Administration (VHA) since April 2017, and a recalibrated version was launched in 2025.

  • baseline

    This shape's defining feature is an absence: the independent model check and the program-evaluation check are drawn empty at baseline, because no independent pre-deployment subgroup or bias validation ran on the model as it scored the whole Veterans Health Administration (VHA) population nationally. A bias assessment was planned but, per the Government Accountability Office's 2022 review, not completed, and the documented demographic miss - women under-flagged because military sexual trauma, intimate-partner violence, and LGBTQ+ identity were excluded - went uncaught until an external investigation and a years-later recalibration (REACH VET 2.0).

  • assumed

    Unlike most tools in this Atlas, the human review here is drawn active, not inactive: the coordinator-to-clinician handoff and the clinician's own discretion are a genuine, staffed, two-stage loop, because the flag is advisory and never triggers an automated care action, and veterans may decline outreach. No public quantitative override or veteran-decline rate exists, so the exact strength is a modest modeling choice.

  • baseline

    The surrogate-versus-target gap is drawn as an inactive reconciliation between two record systems: the electronic health record (EHR), which holds the proximal engagement outcomes the program improved (completed appointments, new safety plans, fewer documented attempts), and a separate mortality repository, which holds the true target outcome the outreach did not move. Two Department of Veterans Affairs (VA) evaluations (2021, 2025) found the proximal improvement without a reduction in death by suicide; the reconciliation that reveals that gap runs only as periodic external study, not routinely.

  • baseline

    The demographic miss is drawn on the store-to-model feature edge, where it is documented to live: a fixed 0.1% budget of enhanced outreach is allocated by predicted risk, so the choice of variables decides who is eligible, and factors specific to women were excluded from the original set. This is why 'improve the model' moves it least - a retraining experiment barely shifted predictive value - and the fix that mattered was changing the variable set (who is eligible), not the accuracy.

  • assumed

    The monthly flag list is drawn as a mediating artifact on the model -> coordinator pathway. Its defining property is that the top-0.1% budget is capped, keeping the list short enough to act on; it carries no flow of its own and does not affect the dynamics.

  • assumed

    No suicide or crisis outcome is modeled here. This Lab models institutional propagation only, and served veterans are not in the dynamics. The documented sex disparity in the served population is recorded externally in the case file, is contested (VA framed the excluded factors as simply less predictive), and is never computed from anything in this diagram.

What this example does not show

  • This Lab models institutional propagation only. It never models suicide or any crisis outcome, and the veterans this flag serves are not in the diagram - a flag or an act of outreach here is an institutional signal, never a life. The improved appointments and the unchanged suicide-death rate that two VA evaluations documented are recorded in the case file, measured outside any diagram like this one.
  • The documented demographic miss - that the model treated being a white man as a stronger risk signal than factors specific to women, and excluded military sexual trauma, intimate-partner violence, and LGBTQ+ identity - is an investigative-journalism finding that VA has contested, framing the excluded factors as simply less predictive. It is carried in the case file as documented-but-contested, and no per-subgroup flag rate is asserted here, because the public subgroup rates are partial and contested.
  • The sensitivity near 2% and false-negative rate near 98% are from an independent re-analysis of 2018 data for the death-by-suicide outcome; do not conflate them with the higher predictive values reported for the combined attempt-or-death outcome. The 2025 recalibration (REACH VET 2.0) had not been independently evaluated for performance or bias as of the research date, so this shape is preserved as the documented 1.0-era structure, not a claim about 2.0.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Veterans Health Administration since 2017, scoring about 6.28 million patients and flagging the top 0.1% at each facility (roughly 6,300 to 6,700 veterans a month, more than 130,000 since 2017); an independent re-analysis of 2018 data found the top-0.1% flag has a positive predictive value near 0.05% and a false-negative rate of about 98% for death by suicide, and a 2024 investigation reported that the model treated being a white man as a stronger risk signal than factors specific to women and excluded military sexual trauma and intimate-partner violence from its variables, a characterization VA has contested by framing the excluded factors as less predictive.

    empirical
    • Academic Harris, Finlay, Meerwijk, Evaluating the accuracy of the VHA REACH VET suicide prediction model for legal involved veterans (npj Mental Health Research, 2025;4:53) https://pmc.ncbi.nlm.nih.gov/articles/PMC12535588/
    • Investigative Glantz, V.A. Uses a Suicide Prevention Algorithm to Decide Who Gets Extra Help. It Favors White Men. (The Markup with The Fuller Project, 2024) https://themarkup.org/news/2024/05/30/v-a-uses-a-suicide-prevention-algorithm-to-decide-who-gets-extra-help-it-favors-white-men
    • Trade press Graham, Inside VA's yearslong AI effort to uncover veterans at high risk of suicide (Nextgov/FCW, 2025) https://www.nextgov.com/artificial-intelligence/2025/07/inside-vas-yearslong-ai-effort-uncover-veterans-high-risk-suicide/406781/
    • Government U.S. Government Accountability Office, Veteran Suicide: VA Efforts to Identify Veterans at Risk through Analysis of Health Record Information (GAO-22-105165, 2022) https://www.gao.gov/assets/gao-22-105165.pdf
  • Two Veterans Health Administration evaluations of REACH VET found the program associated with improved proximal outcomes — more completed outpatient appointments, more new safety plans, and fewer documented suicide attempts — but not with reduced death by suicide: a 2021 triple-differences study of 173,313 veterans across 141 facilities found no association with suicide or all-cause mortality, and a 2025 follow-up of 266,246 observations replicated the null with all confidence intervals crossing one; both are observational rather than randomized studies.

    empirical
    • Academic McCarthy, Cooper, Dent et al., Evaluation of the REACH VET Suicide Risk Modeling Clinical Program in the Veterans Health Administration (JAMA Network Open, 2021;4(10):e2129900) https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2785078
    • Academic Dent, Cooper, McCarthy, The REACH VET Program and Mortality Outcomes Among Veterans at High Risk of Suicide (JAMA Network Open, 2025;8(7):e2519513) https://pmc.ncbi.nlm.nih.gov/articles/PMC12238888/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Behavioral-health & crisis triage domain page.

Levers available here and the patterns behind them

Documented case histories