PAN Lab example
Kaiser Permanente Suicide-Risk Model
The added sensor: an EHR-embedded suicide-risk score
A machine-learning score flags a patient's suicide risk within about 30 minutes of a virtual mental-health intake, and either the score or the existing self-report screen routes that patient into the same assessment-and-outreach workflow. Modeled on Kaiser Permanente's suicide-risk model. Nothing here decides anything: the flag is advisory, the clinician still assesses. Watch what an added, redundant sensor does to a workflow — it can add flags but never remove them, and against a 0.17% base rate almost every one is a false positive.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Kaiser-EHR-class embedded suicide-risk flag network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 6 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the electronic-health-record-embedded, workflow-integrated risk-flag pattern documented in the Kaiser Permanente suicide-risk case file — not a reconstruction of the actual model, dashboards, or workflow.
- baseline
The defining feature is an either-channel merge: the machine-learning flag augments a pre-existing self-report screen — the Patient Health Questionnaire-9 (PHQ-9) and the Columbia-Suicide Severity Rating Scale (C-SSRS) — and either channel triggers the same assessment-and-outreach workflow, so the added sensor can raise the flagged volume but never lower it. That additive-only integration is drawn as two channels landing on one clinician workflow, combined rather than reconciled, with the check that would reconcile them left inactive.
- assumed
Peer pathways are authored on both signs: assessment habits and the weight given the flag spread clinician to clinician, one score homogenizes every intake's blind spots (subgroup discrimination ranged AUROC 0.69 to 0.89 across race groups), while a real, staffed inhibiting check exists — the community-engaged design and clinician-manager iterative testing that reviewed subgroup validity before launch.
- assumed
The combined-but-not-reconciled channels and the absence of standing production monitoring are this shape's defining absence, drawn as an inactive independent model check. The ex-ante subgroup validation was genuine but was not carried into continuous monitoring of who the flag reaches, who acts on it, or how the false-positive load is absorbed; no clinician flag-acceptance, override, assessment-completion, or workload rate is public.
- assumed
The flag is advisory and never gates care; clinician discretion over the assessment and any response is retained throughout, so the correction capacity is genuine. No public quantitative override or acceptance rate exists, so it stays a modest modeling choice.
- assumed
The combined-flag marker sits on the two-channels-to-clinician pathway — it carries no flow of its own and does not affect the dynamics; it marks that the two alert sources are summed rather than reconciled.
- assumed
The measured differences in the model's discrimination across demographic subgroups are recorded in the case file as external observations. This Lab models institutional propagation, not demographics, and estimates no differential harm and no suicide or crisis outcome for served patients.
What this example does not show
- This Lab models institutional propagation only. It never models suicide or any crisis outcome, and the patients this flag serves are not in the diagram — a flag, an assessment, or an act of outreach here is an institutional signal, never a life. The reports on this deployment are feasibility- and design-focused and present no evidence that it reduced suicide attempts; that boundary is recorded in the case file, and any real outcome would be measured outside any diagram like this one.
- The measured differences in the model's discrimination across subgroups (area under the ROC curve from 0.69 to 0.89 across race groups, and 0.79 for women versus 0.74 for men) are recorded in the case file as external observations with wide confidence intervals for small subgroups; no per-subgroup flag rate is asserted here, and the Lab estimates no differential harm to served people.
- The Epic / KP HealthConnect identification of the electronic health record is well-supported context (Kaiser Permanente Northern California runs KP HealthConnect, an Epic system) rather than a direct in-text vendor claim; the peer-reviewed record says electronic health record. No large-language-model or generative component is involved — this is a structured-data penalized-regression risk model, and it should not be read as a generative-AI tool.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Kaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electronic health record of a large virtual mental-health program that handles more than 5,000 intake visits a month; the model is scored in near-real-time (about a 30-minute delay after an encounter trigger) and, at pre-set thresholds, flags high-risk patients to the intake clinician, routing them into the same suicide-risk-assessment and outreach workflow that a positive self-report screen (the PHQ-9 and Columbia-Suicide Severity Rating Scale) triggers, so the machine flag and the self-report alert are effectively OR-merged. In a study of 1,623,232 intake appointments (2012 to 2022, base rate 0.17 percent) the model reached an area under the ROC curve of 0.77 and its top risk decile captured 48.8 percent of appointments later followed by an attempt, but with a positive predictive value of about 0.8 percent.
empirical- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298
- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full
- Academic Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/
Because the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake is very low (0.17 percent) and the positive predictive value in the top risk decile is about 0.8 percent, the large majority of flagged patients will not attempt suicide in the window, so adding the machine-learning flag as a redundant sensor OR-merged onto the existing self-report screen imports a substantial false-positive and clinician-workload burden at scale — a caution the implementation team itself raised. The implementation reports are feasibility- and design-focused and present no evaluation showing the deployment reduced suicide attempts.
empirical- Academic Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/
- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298
- Academic Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Keep skills sharp — Deskilling-arrest mandate
- Escalate checks — State-feedback vigilance
- Check with a second model — Cross-model verification
- Review on schedule — Oversight cadence & retrospectives
- Require sign-off — Conformity assessment gate
- Mark AI-written records — Provenance labeling
- Upgrade model — Improve the model
- Store less data — Data minimization
Documented case histories
- Kaiser Permanente Suicide-Risk Model
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Crisis Text Line & Loris.ai
- LyssnCrisis counselor QA at ProtoCall Services (988)
- NarxCare
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- Tessa chatbot replacing the NEDA eating-disorder helpline
- Woebot (a governed app wind-down)