Skip to content

PAN Lab example

Vanderbilt VSAIL suicide-risk alert

The alert that had to be dismissed: an EHR suicide-risk model

The same suicide-risk score, the same threshold, the same patients — and the only thing that changes is whether the alert interrupts the clinician or waits quietly in the chart. In the record that shape came from, that one change moved point-of-care screening from about 4% to about 42%. Modeled on Vanderbilt's VSAIL suicide-risk alert. The lever here is not the model's accuracy; it is the delivery channel and the attention it commands. Watch the other edge of it: force attention on a low-precision flag too bluntly and you fatigue the very screen you depend on.

Stylized model of a documented deploymentBehavioral-health & crisis triage

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the VSAIL-class EHR suicide-risk alert network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 6 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the interruptive-versus-passive clinical-decision-support pattern documented in the VSAIL / Vanderbilt Safecourse case file — not a reconstruction of the actual model or its alerts.

  • baseline

    The primary channel is the alert itself: the documented finding is that making the identical alert interruptive rather than passive raised point-of-care screening roughly tenfold (about 42% versus 4%) with the underlying score, threshold, and screen held constant. The Lab draws that as the model-to-operator channel starting low — the passive default was ignored far more often than not — and being what an attention lever raises.

  • baseline

    Raising that channel bluntly carries the countervailing risk the researchers named: alert fatigue on a low-precision flag (number-needed-to-screen 271 for attempt in the highest-risk quantile of the 2021 silent-mode study), which erodes the engaged screen the design depends on. The baseline treats the screen as advisory and discretionary — even the interruptive alert produced no screening in about 58% of encounters.

  • assumed

    The alert-rate threshold is drawn as a mediating artifact on the model → alert pathway — the documented, deliberate cutoff that bounded the flag to roughly the top 8% of encounters at about a 2% risk threshold so clinician burden would stay survivable. It carries no flow of its own and does not affect the dynamics; whether the bound was set right is a design question recorded in the case file, not computed here.

  • assumed

    The memory loop from routine electronic health record (EHR) fields into each score is present at baseline: the score is computed from records the patient never provided for this purpose, and race, ZIP-based area deprivation, and mental-health diagnoses are among the features. What those features encode as differential exposure stays external (below).

  • assumed

    The silent-mode monitoring node carries the documented algorithmovigilance loop — the model ran silently and was recalibrated after early miscalibration — drawn as a slow oversight inflow only, not a per-case second read and not a corrective edge back into the model. The recalibration correction is realized through the oversight-cadence lever, not a static edge. It is the only check above a single in-house model.

  • assumed

    Peer pathways are authored on both signs: one model scores every registration (correlated error, worst in the behavioral-health setting where the target is most concentrated) and alert-response habits spread clinician to clinician, while the inhibiting checks are the defining absence — the screen is administered solo with no routine peer second-read, and no independent model re-checks the score, so the only check on a drifting flag is the slow monitoring loop.

  • assumed

    Suicide and crisis outcomes are never modeled here; the Lab models institutional propagation only. Differential exposure to the people the model scores is documented in the case file and measured outside a diagram like this one; the subgroup number-needed-to-screen differences are reported but unadjudicated (bias versus base-rate), and no differential client-harm figure is asserted here.

What this example does not show

  • Suicide and crisis outcomes are never modeled here; the Lab models institutional propagation only. The trial this shape is drawn from measured a process outcome — whether a screen happened — not suicide prevention: no suicidal ideation or attempts were documented in either arm during 30-day follow-up, and the trial was explicitly not powered for clinical outcomes, so nothing here speaks to reduced harm. Differential exposure to the people the model scores is documented in the case file and measured outside any diagram like this one.
  • The headline 42% versus 4% screening figures come from a single-center trial in three ambulatory neurology clinics over six months, not from routine center-wide live alerting; the model is research and operational decision support, not an independently cleared device. The number-needed-to-screen and c-statistic figures come from the separate 2021 silent-mode study, not the trial, and the reported subgroup differences are unadjudicated as bias versus base-rate.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In a single-center randomized trial across three Vanderbilt neurology clinics (August 2022 to February 2023), an EHR suicide-risk model flagged 596 of 7,732 encounters (about 8%) at a 2%-or-higher 30-day-risk threshold; making the identical alert interruptive rather than passive led clinicians to elect a suicide-risk screen in 42% of encounters (121/289) versus 4% (12/307) for a passive chart icon, an adjusted odds ratio of 17.70 (95% CI 6.42–48.79). Screening remained fully advisory: about 58% of interruptive and 96% of passive alerts produced no screening.

    empirical
    • Academic Walsh et al., Risk Model-Guided Clinical Decision Support for Suicide Screening: A Randomized Clinical Trial (JAMA Network Open, 2025; PMC11699529) https://pmc.ncbi.nlm.nih.gov/articles/PMC11699529/
    • Vendor AI tested for alerting clinicians of suicide risk at three VUMC clinics (Vanderbilt University Medical Center News, first-party institutional communication, 2025) https://news.vumc.org/2025/01/03/ai-tested-for-alerting-clinicians-of-suicide-risk-at-three-vumc-clinics/
    • Trade press Suicide prevention more feasible using AI-powered screening alerts (Healio Primary Care, 2025) https://www.healio.com/news/primary-care/20250122/suicide-prevention-more-feasible-using-aipowered-screening-alerts
  • In a separate 2021 prospective silent-mode study (115,905 predictions on 77,973 patients, June 2019 to April 2020), the model reported a c-statistic of 0.797 for suicide attempt and 0.836 for ideation center-wide but only 0.544 for attempt in behavioral-health settings, and in the highest-risk quantile the number-needed-to-screen was 271 for attempt and 23 for ideation. In the 2022 to 2023 trial no suicidal ideation or attempts were documented in either arm during 30-day follow-up, and the trial was explicitly not powered for clinical outcomes, so it measured a process outcome (screening) rather than reduced harm.

    empirical
    • Academic Walsh et al., Prospective Validation of an Electronic Health Record-Based, Real-Time Suicide Risk Model (JAMA Network Open, 2021; PMC7955273) https://pmc.ncbi.nlm.nih.gov/articles/PMC7955273/
    • Academic Walsh et al., Risk Model-Guided Clinical Decision Support for Suicide Screening: A Randomized Clinical Trial (JAMA Network Open, 2025; PMC11699529) https://pmc.ncbi.nlm.nih.gov/articles/PMC11699529/
    • Trade press Suicide prevention more feasible using AI-powered screening alerts (Healio Primary Care, 2025) https://www.healio.com/news/primary-care/20250122/suicide-prevention-more-feasible-using-aipowered-screening-alerts

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Behavioral-health & crisis triage domain page.

Levers available here and the patterns behind them

Documented case histories