Skip to content

PAN Lab example

Stratification Tool for Opioid Risk Mitigation

The health system that randomized its own oversight lever

A national health system scores every patient with an opioid prescription for one-year risk of an overdose-related or suicide-related event, refreshes the score nightly, and requires an interdisciplinary team at all 140 of its medical centers to case-review whoever lands in the top percentile. Modeled on a documented deployment, and on the rarest evidence in this atlas: the system randomized its own governance lever. It randomized when facilities widened the mandated tier from the top 1% to the top 5%, and separately whether the policy notice carried an accountability paragraph - a 97% completion target, quarterly reporting, technical assistance and action plans. The mandate worked on process: flagged patients became 5.1 times more likely to be reviewed. The accountability wrapper ran the other way - the facilities that got it completed fewer reviews. Watch what the oversight edge does when you lean on it. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. Cost is not what blocks it. Inside the budget the best legal settings clear the benefit margins on both tiers, and lifting the pathway requirement on its own wins both tiers from a stack costing 4 of the 13 you have. The pathway gate is the only gate that fails. Six pathways stay open at every affordable price, and together they are the mandate. The warehouse feeds the nightly score. The care record is read back into the next night's run. The compelled review and the prescribing both write into that record, the completion count is built from it, and the trial reads its outcomes out of it. Closing those would mean switching the programme off. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.

Stylized model of a documented deploymentBehavioral-health & crisis triage

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Percentile-mandate-class opioid risk review network: 11 components and 20 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 1 published baseline · 4 measured. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    Topology. Eleven components, all documented, none decorative. The three reviewers are the three oversight layers the record enumerates, and each has a different documented inbound read: the national policy office reads a quarterly compliance figure, the randomized evaluation partnership reads the patient record at full population scale (44,042 / 16,272 / 10,685 / 28,251 patients across 140 facilities and 23 months), and the government audit office reads a 103-record sample at 5 of 140 facilities. Two absences are equally derived: there is no enforcement component, because the mandate compels a review and never an action - no mandated taper, no automatic prescription cutoff tied to the score - and there is no staff-to-model pathway, because the score is computed from dispensing records, diagnoses, prior events and utilization history, so no clinician's framing can bend it. No external boundary either: the deployment is built in-house on the department's own warehouse with no commercial vendor and no data broker, and the record documents no egress of patient data.

  • measured

    Heavy workload against limited capacity. Workload: a review mandate attached to a percentile of a score that recomputes nightly across 140 medical centers, with the mandated cohort deliberately widened from the top 1% to the top 5% at two dated waves. Capacity above the very-low floor: the human channel is a separately-mandated clinical function - pain-management teams required by addiction-and-recovery legislation independent of this tool, clinical pharmacists, mental-health and suicide-prevention staff, academic detailing and pain champions - and opioid case review existed before the mandate at a 6.6% rate. Capacity held below full: median facility completion 71% with an interquartile range of 48% to 95%, about one facility in five at the 97% target, roughly 57% of very-high-risk patients reviewed, and an audit record sample in which 21 of 53 long-term-opioid-therapy patients had no urine drug screening in the prior year, 40 had no monitoring-programme query and 12 had no written informed consent.

  • measured

    Baselines. The two record-to-model pathways run at full strength because the whole predictor set is recomputed nightly for a population the development cohort alone put at 1,135,601 patients, and because the record-to-model loop closes in one day - the shortest documented record-to-model cycle in this atlas. The model-to-staff pathway into the review teams is at full strength because national policy compels it and it produced the largest measured behaviour change in the record (case-review odds ratio 5.1, 95% confidence interval 3.64 to 7.23, from a pre-mandate rate of 6.6%); into prescribers it is at a substantial level because facility-level dashboard use was reported by 97% of the 89 facilities surveyed while no source reports a per-clinician action rate and no clinical action is compelled. The evaluation read is at full strength and the audit read faint because one measured the whole population for 23 months and the other sampled 103 records at 5 facilities for one year. The accountability check is faint because the department randomized it and measured it running against its own aim.

  • measured

    The percentile gate as a model-to-operator check, and the honest form of that pathway: a bounded automated screen whose partiality is published rather than inferred. The tier cut is applied automatically to every nightly score and decides which of them becomes a compelled review - a score of 0.166 or above for the top 1% tier, later the top 5% - and the deployment's own figures state its reach: the highest-risk 8.8% of the scored population was estimated to capture 50.2% of events, with a 7.9% false-positive figure. It bounds the mandate's span and is never a substitute for reading the patient. It is drawn at a substantial level, not full, because that published reach is about half the events. The threshold values are time-indexed: the top-1% configuration and the 0.166 score belong to the period before the stepped-wedge waves of 12 February 2019 and 13 August 2019.

  • baseline

    The latent second read. The independent model check is drawn empty on a documented absence, per the derivation rule that a check drawn present at any strength requires a source documenting that it exists. The independent critique of this algorithm class holds this tool up as its better-evidenced example - full-population technical work plus a randomized clinical evaluation - while recording that no external validation of the class has been published and asking for subgroup analyses and independent oversight as well. The successor models the department has published were built on the same administrative data, which is why the monoculture pressure and the second-read lever both bear on this deployment.

  • measured

    The benefit and the harm travel together, and this diagram must never be read as saying otherwise. The four-month all-cause mortality odds ratio of 0.78 (95% confidence interval 0.65 to 0.94) is an exploratory endpoint; the pre-specified primary composite of nine serious-adverse-event categories did not move. Among 10,685 high-risk long-term opioid therapy patients the mandate cut opioid discontinuation by 11.16 percentage points and all-cause mortality by 3.31 percentage points against baselines of 29.1% and 9.5%. Among patients newly diagnosed with opioid use disorder during the trial, 90-day mortality odds were 1.74 (95% confidence interval 1.06 to 2.87) with no significant change in serious adverse events. The two randomizations are distinct experiments: the stepped-wedge threshold expansion and the accountability-language arm. The subgroup findings ride the first; the oversight-backfire finding belongs to the second alone.

  • assumed

    Evidence status. Nearly every quantitative figure on this diagram comes from a department-affiliated research-operations partnership or from the national audit office. That partnership published its own null primary composite, its own accountability backfire and its own subgroup harm signal, which is why the evidence is strong enough to draw on; it is peer-reviewed agency self-evaluation and it is not third-party replication, and no independent journalism carrying figures on this deployment was found. The policy notice itself was not retrieved directly - its number, date and content come from the peer-reviewed papers describing it. The targeting concentration figure reaches this bundle through a federal patient-safety profile summarising the development paper. Continued operation rests on that profile (2023) and the audit office's December 2025 status report; the 2024 pharmacist-led facility report is a retrospective chart review of 17 patients reviewed January-September 2018 and evidences the interdisciplinary-team model only.

  • assumed

    No veteran outcome is computed here. This Lab reads institutional propagation only: scores, tiers, reviews, completion counts and reporting cycles. Overdoses, suicide-related events and deaths are served-person outcomes documented in the case file from the trial and audit record, and they are never derived from anything on this diagram. Differential effects across subpopulations - the long-term-opioid-therapy benefit and the newly-diagnosed-OUD harm signal - are carried as documented findings and are never computed from the dynamics.

What this example does not show

  • No veteran outcome is modeled. The Lab reads institutional propagation only - scores, tiers, reviews, completion counts and reporting cycles. Overdoses, suicide-related events and deaths are served-person outcomes documented in the case file from the trial and audit record, and they are never computed from anything on this diagram.
  • The benefit and the harm travel together and neither may be shown alone. The four-month all-cause mortality odds ratio of 0.78 is an EXPLORATORY endpoint of a trial whose pre-specified primary composite of nine serious-adverse-event categories did not move. Among 10,685 high-risk long-term opioid therapy patients the mandate cut discontinuation by 11.16 percentage points and all-cause mortality by 3.31. Among patients NEWLY DIAGNOSED with opioid use disorder during the trial, 90-day mortality odds were 1.74. Nothing here presents the mortality benefit as the trial's confirmed primary result.
  • Two distinct randomizations run in this record and they never merge: the stepped-wedge expansion of the mandated tier from the top 1% to the top 5%, executed 12 February 2019 and 13 August 2019, and the accountability-language arm across 70 facilities against 70. The oversight-backfire finding belongs to the second experiment alone; the subgroup findings ride the first.
  • Nearly every quantitative figure comes from a department-affiliated research-operations partnership or from the national audit office. That partnership published its own null, backfire and harm findings, which is why the evidence is strong enough to draw on - and it is peer-reviewed agency self-evaluation, not third-party replication. No independent journalism carrying figures on this deployment was found, and the policy notice itself was not retrieved directly: its number, date and content come from the peer-reviewed papers describing it.
  • Thresholds and population definitions are time-indexed and the diagram fixes one reading of them. The top-1% tier at a score of 0.166 belongs to the period before the expansion waves; the scored population widened over time to include patients treated for opioid use disorder in the past year. Continued operation rests on a 2023 federal patient-safety profile and a December 2025 audit status report - the 2024 pharmacist-led facility report is a retrospective chart review of 17 patients reviewed January-September 2018 and evidences the interdisciplinary-team model only.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The Veterans Health Administration built an opioid overdose and suicide risk model in-house, with no commercial vendor, on its own electronic health record and Corporate Data Warehouse: fitted on 1,135,601 patients with an opioid prescription in fiscal 2010 against 23,790 overdose-related or suicide-related events among them in fiscal 2011 (a 2.1% base rate), reporting an area under the curve above 0.80 in training and test sets, and refreshed nightly as a continuous one-year risk estimate binned into percentile tiers on a population-management dashboard that shows each patient's risk factors and the guideline-recommended mitigation actions rather than a bare number. Its predictors are administrative - demographics, pharmacy records including opioid type and dose and co-prescribed sedatives, mental health and substance use disorder diagnoses, prior overdose-related and suicide-related events, detoxification episodes, and emergency department and other utilization history - so no structured risk questionnaire and no clinician-scored instrument feeds the score. VHA Notice 2018-08, issued 8 March 2018 and effective 18 April 2018, required all 140 VHA medical centers to convene interdisciplinary teams and case-review every patient in the very-high-risk tier, initially the top 1% of scores at a threshold of 0.166. The mandate compels an assessment of risk, prescription appropriateness, and mitigation options and never an action: there is no mandated taper and no automatic prescription cutoff tied to the score, and every clinical decision remains with the treatment team. Across 44,042 patients in the top 1% to 5% band from April 2018 to March 2020 the mandate raised the odds of receiving a case review 5.1-fold (95% CI 3.64-7.23) and added 0.498 risk-mitigation strategies per patient (95% CI 0.39-0.61). Facility completion had a median of 71% (interquartile range 48-95%), about one facility in five met the 97% target, and each of the 89 surveyed facilities used a median of 23 distinct implementation strategies. Team composition and workflow were left to each facility.

    empirical
    • Academic Oliva, Bowe, Tavakoli et al., Development and applications of the Veterans Health Administration's Stratification Tool for Opioid Risk Mitigation (STORM) to improve opioid safety and prevent overdose and suicide (Psychological Services, 2017;14(1):34-49; doi 10.1037/ser0000099) https://pubmed.ncbi.nlm.nih.gov/28134555/
    • Academic Strombotne, Legler, Minegishi, Trafton, Oliva, Lewis, Sohoni, Garrido, Pizer, Frakt, Effect of a Predictive Analytics-Targeted Program in Patients on Opioids: a Stepped-Wedge Cluster Randomized Controlled Trial (Journal of General Internal Medicine, 2023;38(2):375-381; doi 10.1007/s11606-022-07617-y) https://pubmed.ncbi.nlm.nih.gov/35501628/
    • Academic Minegishi, Garrido, Lewis, Oliva, Pizer, Strombotne, Trafton, Tenso, Sohoni, Frakt, Randomized Policy Evaluation of the Veterans Health Administration Stratification Tool for Opioid Risk Mitigation (STORM) (Journal of General Internal Medicine, 2022; doi 10.1007/s11606-022-07622-1) https://pmc.ncbi.nlm.nih.gov/articles/PMC9585134/
    • Academic Rogal, Chinman, Gellad, Mor, Zhang, McCarthy, Mauro, Hale, Lewis, Oliva, Trafton, Yakovchenko, Gordon, Hausmann, Tracking implementation strategies in the randomized rollout of a Veterans Affairs national opioid risk management initiative (Implementation Science, 2020;15:48; doi 10.1186/s13012-020-01005-y) https://pmc.ncbi.nlm.nih.gov/articles/PMC7313133/
  • The Veterans Health Administration randomized two features of its own opioid case-review policy across its 140 medical centers, and the two randomizations are separate experiments whose findings do not combine. The first was timing: when a facility's mandated review tier widened from the top 1% of risk scores to the top 5%, executed as a stepped wedge with waves on 12 February 2019 and 13 August 2019, which the primary trial paper places at study months 11 and 17. The second was language: whether a facility's copy of the policy notice carried an accountability paragraph naming a 97% case-review completion target, with quarterly reporting to the national Office of Mental Health and Suicide Prevention and technical assistance and action plans for facilities below it. Seventy facilities received that paragraph and seventy did not. Across 16,272 very-high-risk patients (8,734 in the accountability arm, 7,538 outside it), about 57% received a case review overall against a pre-mandate baseline of 6.6%, and there was no difference between arms in opioid-related serious adverse events (hazard ratio 1.03, 95% CI 0.97-1.08) or mortality (hazard ratio 1.00, 95% CI 0.91-1.09) - but patients at accountability-arm facilities were less likely to receive a case review at all (hazard ratio 0.91, 95% CI 0.87-0.95). The implementation tracking study reports the same result at facility level: median completion 71% (interquartile range 48-95%), 18 of 89 surveyed facilities (20%) meeting the 97% target, and facilities given the plain mandate meeting it more often than facilities given the accountability language, 30% against 11% (p=0.04). The practice associated with higher completion was regular self-monitoring and adaptation inside the facility (adjusted incidence rate ratio 1.40) rather than reporting upward from it; dashboard use was reported by 97% of surveyed facilities and local opinion leaders by 80%, while patient-engagement strategies were used by 13%. The oversight-backfire finding belongs to the language experiment alone.

    empirical
    • Academic Minegishi, Garrido, Lewis, Oliva, Pizer, Strombotne, Trafton, Tenso, Sohoni, Frakt, Randomized Policy Evaluation of the Veterans Health Administration Stratification Tool for Opioid Risk Mitigation (STORM) (Journal of General Internal Medicine, 2022; doi 10.1007/s11606-022-07622-1) https://pmc.ncbi.nlm.nih.gov/articles/PMC9585134/
    • Academic Strombotne, Legler, Minegishi, Trafton, Oliva, Lewis, Sohoni, Garrido, Pizer, Frakt, Effect of a Predictive Analytics-Targeted Program in Patients on Opioids: a Stepped-Wedge Cluster Randomized Controlled Trial (Journal of General Internal Medicine, 2023;38(2):375-381; doi 10.1007/s11606-022-07617-y) https://pubmed.ncbi.nlm.nih.gov/35501628/
    • Academic Rogal, Chinman, Gellad, Mor, Zhang, McCarthy, Mauro, Hale, Lewis, Oliva, Trafton, Yakovchenko, Gordon, Hausmann, Tracking implementation strategies in the randomized rollout of a Veterans Affairs national opioid risk management initiative (Implementation Science, 2020;15:48; doi 10.1186/s13012-020-01005-y) https://pmc.ncbi.nlm.nih.gov/articles/PMC7313133/
  • The benefit and the harm signals from the Veterans Health Administration's mandated opioid case-review policy travel together and neither may be reported alone. The widely quoted mortality result - four-month all-cause mortality odds of 0.78 (95% CI 0.65-0.94) - is an EXPLORATORY endpoint of a trial whose pre-specified primary composite of nine serious-adverse-event categories did not move (odds ratio 0.995, 95% CI 0.875-1.132). Among patients NEWLY DIAGNOSED with opioid use disorder during the trial (28,251 analyzed, estimated off the stepped-wedge threshold-expansion randomization rather than the accountability-language arm) the mandate was associated with 90-day all-cause mortality odds of 1.74 (95% CI 1.06-2.87) with no significant change in serious adverse events; a post-hoc subgroup with an opioid prescription before but not after diagnosis showed 5.87 (95% CI 1.85-18.58). Nearly every quantitative source on this deployment is a department-affiliated research-operations partnership: this is peer-reviewed agency self-evaluation that published its own null, backfire, and harm findings, and it is not third-party replication.

    empirical
    • Academic Strombotne, Legler, Minegishi, Trafton, Oliva, Lewis, Sohoni, Garrido, Pizer, Frakt, Effect of a Predictive Analytics-Targeted Program in Patients on Opioids: a Stepped-Wedge Cluster Randomized Controlled Trial (Journal of General Internal Medicine, 2023;38(2):375-381; doi 10.1007/s11606-022-07617-y) https://pubmed.ncbi.nlm.nih.gov/35501628/
    • Academic Auty, Barr, Frakt, Garrido, Strombotne, Effect of a Veterans Health Administration mandate to case review patients with opioid prescriptions on mortality among patients with opioid use disorder: a secondary analysis of the STORM randomized control trial (Addiction, 2023;118(5):870-879; doi 10.1111/add.16110) https://pubmed.ncbi.nlm.nih.gov/36495477/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Behavioral-health & crisis triage domain page.

Levers available here and the patterns behind them

Documented case histories