Domain Atlas / Behavioral-health & crisis triage
Stratification Tool for Opioid Risk Mitigation
The Veterans Health Administration built an opioid overdose and suicide risk model in-house, with no commercial vendor, on its own electronic health record and Corporate Data Warehouse: fitted on 1,135,601 patients with an opioid prescription in fiscal 2010 against 23,790 overdose-related or suicide-related events among them in fiscal 2011 (a 2.1% base rate), reporting an area under the curve above 0.80 in training and test sets, and refreshed nightly as a continuous one-year risk estimate binned into percentile tiers on a population-management dashboard that shows each patient's risk factors and the guideline-recommended mitigation actions rather than a bare number. Its predictors are administrative - demographics, pharmacy records including opioid type and dose and co-prescribed sedatives, mental health and substance use disorder diagnoses, prior overdose-related and suicide-related events, detoxification episodes, and emergency department and other utilization history - so no structured risk questionnaire and no clinician-scored instrument feeds the score. VHA Notice 2018-08, issued 8 March 2018 and effective 18 April 2018, required all 140 VHA medical centers to convene interdisciplinary teams and case-review every patient in the very-high-risk tier, initially the top 1% of scores at a threshold of 0.166. The mandate compels an assessment of risk, prescription appropriateness and mitigation options and never an action: there is no mandated taper and no automatic prescription cutoff tied to the score, and every clinical decision remains with the treatment team. Across 44,042 patients in the top 1% to 5% band from April 2018 to March 2020 the mandate raised the odds of receiving a case review 5.1-fold (95% CI 3.64-7.23) and added 0.498 risk-mitigation strategies per patient (95% CI 0.39-0.61). Facility completion had a median of 71% (interquartile range 48-95%), about one facility in five met the 97% target, and each of the 89 surveyed facilities used a median of 23 distinct implementation strategies. Team composition and workflow were left to each facility.[4]
What happened
The Stratification Tool for Opioid Risk Mitigation is a risk score the Veterans Health Administration built for itself. There is no commercial vendor: it was developed in-house by VHA's Program Evaluation and Resource Center and its Office of Mental Health and Suicide Prevention, on the department's own electronic health record and Corporate Data Warehouse. The model was fitted on 1,135,601 VHA patients who had an opioid prescription in fiscal 2010 and the 23,790 overdose-related or suicide-related events recorded among them in fiscal 2011, a 2.1% base rate, and reported an area under the curve above 0.80 in both training and test sets, a measure of how well the score ranks a patient who went on to have an event above one who did not. Its predictors are administrative: demographics, pharmacy records carrying opioid type and dose and co-prescribed sedatives such as benzodiazepines, mental health and substance use disorder diagnoses, prior overdose-related and suicide-related events, detoxification episodes, and emergency department and other utilization history. No structured risk questionnaire and no clinician-scored instrument feeds it: the inputs are administrative records only. The output is a continuous one-year risk estimate binned into percentile tiers on a population-management dashboard, refreshed nightly, and the dashboard shows each patient's own risk factors and the guideline-recommended mitigation actions available - naloxone, urine drug screening, prescription-monitoring-programme checks, medication for opioid use disorder, mental-health engagement - rather than a bare number.
On 8 March 2018 VHA issued Notice 2018-08, effective 18 April 2018, requiring every medical center to convene an interdisciplinary team and case-review every patient the score placed in the very-high-risk tier - initially the top 1% of scores, at a threshold of 0.166. What the teams must do is assess: the patient's risk, the appropriateness of the prescription, and the mitigation options available. What they are not required to do is act. There is no mandated taper and no automatic prescription cutoff tied to the score; every clinical decision stays with the treatment team. Team composition and workflow were left to each of the 140 facilities, and typically draw on pain-management teams required separately by addiction-and-recovery legislation, clinical pharmacists, and mental-health and suicide-prevention staff. A brief report from one facility describes a pharmacist-led very-high-risk interdisciplinary team; it is a retrospective chart review of 17 patients reviewed between January and September 2018 and evidences that model of team, not later operation.
Then the unusual part. VHA randomized two features of its own policy across the 140 centers and ran a formal evaluation with a research-and-clinical-operations partnership between the Partnered Evidence-based Policy Resource Center, the Center for Health Equity Research and Promotion, the Office of Mental Health and Suicide Prevention and a Veterans Engineering Resource Center. The first randomization was timing: when a facility's mandated tier widened from the top 1% to the top 5%, executed as a stepped wedge with waves on 12 February 2019 and 13 August 2019, which the primary paper places at study months 11 and 17. The second was language: whether a facility's copy of the notice carried an extra accountability paragraph naming a 97% case-review completion target, with quarterly reporting to the national office and technical assistance and action plans for facilities below it. Seventy facilities received that paragraph and seventy did not. These are two distinct experiments and their findings do not combine.
The timing experiment, reported over 44,042 patients in the top 1% to 5% risk band (32,197 control person-periods against 11,845 treatment) from April 2018 to March 2020, found the mandate raised the odds of a patient receiving a case review 5.1-fold, 95% confidence interval 3.64 to 7.23, and added about half a risk-mitigation strategy per patient (0.498, 95% confidence interval 0.39 to 0.61). Exploratory four-month all-cause mortality odds were 0.78, 95% confidence interval 0.65 to 0.94. The pre-specified primary composite of nine serious-adverse-event categories did not move (odds ratio 0.995, 95% confidence interval 0.875 to 1.132). The mortality figure is exploratory and the primary endpoint was null; both facts belong to any statement about what this mandate achieved.
The language experiment, reported over 16,272 very-high-risk patients (8,734 at accountability-arm facilities, 7,538 at the others), found the opposite of its design intent. About 57% of very-high-risk patients received a case review overall, against a pre-mandate baseline of 6.6%. Between arms, opioid-related serious adverse events were unchanged (hazard ratio 1.03, 95% confidence interval 0.97 to 1.08) and so was mortality (hazard ratio 1.00, 0.91 to 1.09) - but patients at the accountability facilities were less likely to receive a case review at all, hazard ratio 0.91, 95% confidence interval 0.87 to 0.95. The implementation tracking study puts facility numbers on the same result: median facility case-review completion of 71% with an interquartile range of 48% to 95%, about one facility in five meeting the 97% target, and facilities given the plain mandate meeting it more often than facilities given the accountability language, 30% against 11% (p=0.04). Facilities improvised heavily around the mandate: of the 89 facilities surveyed, each used a median of 23 distinct implementation strategies (interquartile range 16 to 31), with dashboard use at 97% and local opinion leaders at 80% while patient-engagement strategies were used by 13%; regular self-monitoring and adaptation was associated with higher completion, adjusted incidence rate ratio 1.40.
One secondary analysis cuts against the trial's exploratory benefit, and it belongs to the record. Among patients newly diagnosed with opioid use disorder during the trial - 28,251 analyzed, and the estimate rides the stepped-wedge threshold-expansion randomization rather than the language arm - the mandate was associated with 90-day all-cause mortality odds of 1.74, 95% confidence interval 1.06 to 2.87, with no significant change in serious adverse events; a post-hoc subgroup with an opioid prescription before but not after diagnosis showed 5.87 (1.85 to 18.58).
The sociotechnical reading
The structural fact worth carrying out of this case is that the governance lever was the treatment. Everywhere else in this atlas the oversight arrangement is a given: it exists, someone describes it, and whether it works has to be argued from outcomes that were never randomized against anything. Here a national health system took its own accountability mechanism - a completion target, a quarterly report, technical assistance, action plans - wrote it into half the copies of its own policy notice, withheld it from the other half, and measured what happened. The answer is that the facilities that received the pressure completed fewer reviews than the facilities that did not, with target attainment of 11% against 30%, and that the difference showed up in review receipt and nowhere in adverse events or mortality. That is an empirical datum about oversight-pressure dynamics, and it is nearly unique. It also has a mechanism the implementation record points at: the practice associated with higher completion was regular self-monitoring and adaptation inside the facility, not reporting upward from it. A number that travels up a hierarchy on a quarterly cycle is a different instrument from a number a team checks against its own work on Tuesday, and this deployment is the closest thing anyone has to an experiment separating them.
The second reading is about what a mandate on a moving frontier costs. The obligation attaches to a percentile of a score that recomputes every night, so the population owed a review changes daily and the threshold that defines it is a dial the system can turn - which it did, from the top 1% to the top 5%, on dates it randomized. Widening a tier is arithmetic on the demand side, and nothing in the notice added anything on the supply side: no staffing came with it. What the trial measured is that the newly covered patients did get reviewed more, their odds of a case review rising 5.1-fold. What the implementation record shows alongside that is that completion stayed partial as the covered population grew - a median facility completion of 71% with an interquartile range running from 48% to 95%, and roughly 57% of very-high-risk patients reviewed. Meanwhile the mandate's own writes flow back into its inputs faster than anywhere else in this atlas. A completed review documents an assessment and adds mitigation strategies into the same record the score reads at the next nightly refresh, so the intervention and its own input sit one day apart. And what reaches the national office is not the record but a quarterly count derived from it - the compliance surface, not the clinical one, which is exactly why an accountability paragraph aimed at the count could move the count's denominator without moving care.
The third reading is the hardest and the one this case exists to hold open: the same lever helped and may have harmed, inside one randomized frame, and the deploying system published both. Among patients newly diagnosed with opioid use disorder during the trial, 90-day mortality odds were 1.74. The headline mortality benefit that gets quoted, an odds ratio of 0.78, is an exploratory endpoint from a trial whose pre-specified primary composite did not move. None of these can be reported without the others, and the deployment's own evaluators are the reason all of them are in the public record - which is the credit and the caveat in one sentence, because nearly every quantitative source here is department-affiliated, and no outside party has replicated any of it. What lands on veterans - an overdose, a suicide attempt, a death, a prescription continued or stopped - is documented in the trial and audit evidence above and is never computed from any diagram.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library.