Domain Atlas / Clinical decision support & deterioration alerting
UBH Level of Care Guidelines (Wit v. UBH)
Explore this deployment in the PAN Lab ↗
In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.
The least-established input behind this case's model organization's readings is an assumption, not a measurement. Evidence base: 1 assumed · 8 published baseline.
After a ten-day ERISA bench trial in October 2017, the United States District Court for the Northern District of California issued 106 pages of Findings of Fact and Conclusions of Law (28 February 2019; public redacted version 5 March 2019) holding that the 2011-2017 editions of United Behavioral Health's Level of Care Guidelines and Coverage Determination Guidelines were significantly and pervasively more restrictive than generally accepted standards of care. The operative 3 February 2026 judgment declares eight specific deviations, among them excessive emphasis on acute crisis stabilization, no effective treatment of co-occurring conditions, failure to err toward a higher level of care where the indicated level is ambiguous, no coverage to maintain function, motivation-based exclusions in the 2014-2017 editions, no child-and-adolescent-specific criteria, an overbroad custodial-care exclusion paired with a narrow active-treatment requirement, and mandatory prerequisites in place of a multidimensional assessment; the editions also omitted the ASAM residential levels 3.1, 3.3 and 3.5. The court found the criteria operated as binding rules rather than guidance: only a physician or doctoral-level psychologist could issue a clinical non-coverage determination, such a reviewer typically spent about thirty minutes talking to the requesting physician and writing up conclusions, every denial letter had to cite the specific guideline relied on, and the company's testimony that reviewers could deviate from the guidelines on clinical judgment was found not credible. No machine-learning system is involved; the guidelines are codified decision criteria applied by human reviewers.[2]
What happened
The algorithm in this case is a document. United Behavioral Health — operating as OptumHealth Behavioral Solutions, a UnitedHealth Group company, and acting as an Employee Retirement Income Security Act (ERISA) claims administrator and fiduciary for commercial welfare benefit plans — wrote its own Level of Care Guidelines and Coverage Determination Guidelines: codified admission, continued-stay and discharge criteria for each level of mental-health and substance-use care. It reviewed and reissued them at least once a year from 2011 through 2017, distributed them to every reviewer, and kept them uniform across fully-insured plans, where it carries the benefit-expense risk itself, and self-funded plans, where it does not. There is no machine-learning system anywhere in this record, and the court's findings are not about an error rate. They are about who was in the room when the criteria were written.
The pipeline was short and its authority was sharply divided. A member or provider requests coverage. A Care Advocate — a front-line clinician employed by the company — collects the member's clinical information and applies the guidelines. A Care Advocate may approve on clinical grounds and may deny on administrative ones, but a clinical denial must be passed to a Peer Reviewer: a physician or doctoral-level psychologist, the only class authorized to issue a Clinical Non-Coverage Determination. The trial record puts that review at approximately thirty minutes talking to the requesting physician and writing up conclusions, plus additional time on the Care Advocate's file. The company's own 2014 Utilization Management Program Description required every denial letter to state the rationale and cite the specific guideline relied on — a requirement that, years later, is what let a court define the classes, because every class member's denial letter or case note named the Guidelines. Company witnesses testified that Peer Reviewers could deviate from the criteria on clinical judgment. The court found that testimony not credible.
Upstream sat the part the trial made visible. Drafting workgroups wrote and revised each year's edition and submitted the revisions for approval to the Behavioral Policy and Analytics Committee, from 2011 to 2016, and then to its successor the Utilization Management Committee. Named representatives of the company's Finance and Affordability Departments sat on those committees throughout the class period. The court's finding, at paragraph 180 of its Findings of Fact, is stated plainly: "The Court finds that the financial incentives discussed above have, in fact, infected the Guideline development process. In particular, instead of insulating its Guideline developers from these financial pressures, UBH has placed representatives of its Finance and Affordability Departments in key roles in the Guidelines development process throughout the class period." Proposed changes were modelled for their impact on benefit expense before adoption; the record shows proposed liberalizations of the criteria waiting on a "green light" from finance. The company prepared detailed benefit-expense forecasts and targets, tracked monthly trends, and took action to address benefit expenses that exceeded its projections. A 2014 internal presentation named continued use of concurrent review to ensure appropriate utilization as the "Mitigation Strateg[y]" for the 2008 Parity Act's removal of day and visit limits. Because the criteria were kept uniform across both plan categories, the court held the conflict tainted the company's decision-making as to both.
After the ten-day bench trial in October 2017, the court issued 106 pages of Findings of Fact and Conclusions of Law on 28 February 2019, with the public redacted version filed 5 March 2019. It found the 2011-2017 editions significantly and pervasively more restrictive than generally accepted standards of care. The operative 2026 judgment enumerates eight deviations: excessive emphasis on acute crisis stabilization rather than the underlying condition; no effective treatment of co-occurring conditions; failure to err in favour of a higher level of care where the indicated level is ambiguous; no coverage of treatment to maintain function or prevent deterioration; motivation-based exclusions in the 2014-2017 editions; no child-and-adolescent-specific criteria; an overbroad custodial-care exclusion paired with a narrow active-treatment-and-improvement requirement; and mandatory prerequisites in place of a multidimensional assessment of the individual. The editions also omitted the American Society of Addiction Medicine (ASAM) residential levels 3.1, 3.3 and 3.5. On the company's own witnesses the court wrote: "UBH's experts, on the other hand, had serious credibility problems. The Court found that with respect to a significant portion of their testimony each of them was evasive — and even deceptive — in their answers when confronted with contrary evidence."
Four states had already required something else. Connecticut mandated ASAM criteria for these determinations from 1 October 2013, Illinois from 18 August 2011, Rhode Island from 10 July 2015, and Texas required the criteria of its Department of Insurance. The court adjudicated violations of all four. It also found that the company's 2013 and 2015 "crosswalks" told Connecticut regulators that all three ASAM residential levels were included in its admission criteria and that, at the time these statements were made to Connecticut regulators, the company knew them to be false. United Behavioral Health did not appeal that portion of the judgment, and the Ninth Circuit recorded that it therefore remains intact — the one part of the case that came through every appellate cycle unchanged.
The rest of the case did not. On 3 November 2020 the district court ordered sweeping relief: reprocessing of approximately 67,000 coverage requests, a ten-year injunction with a five-year modification check-in requiring criteria consistent with generally accepted standards — naming ASAM, the LOCUS assessment scale, the Child and Adolescent Service Intensity Instrument (CASII) and the Early Childhood Service Intensity Instrument (ECSII) absent contrary state law — a court-appointed special master, and a supervised retraining programme for Care Advocates, Peer Reviewers and external clinical consultants with sixty- and ninety-day completion gates and an annual refresh. The Ninth Circuit stayed reprocessing on 12 February 2021. That order never took effect. In an unpublished memorandum of 22 March 2022 (later withdrawn), a published opinion of 26 January 2023 (also withdrawn), and finally the amended opinion of 22 August 2023 — Wit III, 79 F.4th 1068 — the Ninth Circuit affirmed Article III standing and affirmed certification of the three classes for the fiduciary-duty claim, but reversed certification of the denial-of-benefits classes under the Rules Enabling Act, held that the district court erred to the extent it determined that the ERISA plans required the guidelines to be coextensive with generally accepted standards of care, held that reprocessing was not appropriate equitable relief under 29 U.S.C. section 1132(a)(3), and remanded the exhaustion question. When the district court's scope-of-remand order tried to preserve more than the mandate allowed, the company obtained mandamus: on 4 September 2024, in No. 24-242, the Ninth Circuit directed entry of judgment for United Behavioral Health on the denial-of-benefits claim, writing that in its thorough analysis of the spirit of the mandate, the district court lost the letter.
What survived, and what it produced, is narrower and more specific than the 2020 order. On 5 August 2025 the district court held that the fiduciary claim survives Wit III insofar as it rests on the duties of loyalty and due care, entered judgment for the company on the duty-to-follow-plan-terms theory, and held that the surviving statutory claim requires no administrative exhaustion — alternatively, that exhaustion is excused as futile on the trial findings. The Amended Remedies Order of 3 February 2026 then vacated the November 2020 order in its entirety and superseded it. It declares that the misconduct in developing and adopting the guidelines was willful and systematic, that the adjudicated editions are irreparably tainted by the company's disloyalty and lack of care, and that the company breached 29 U.S.C. sections 1104(a)(1)(A) and (B) and violated the four state mandates. It permanently enjoins use of those editions to implement plan terms about generally accepted standards of care. And it orders that for five years, through 3 February 2031, with jurisdiction retained, any criteria the company adopts for that purpose shall accurately reflect those standards as established in the court's Findings of Fact and the requirements of any applicable state law. There is no special master in it, no retraining programme, no court-specified catalogue of external criteria, and no reprocessing. No coverage request was ever reprocessed. Attorney-fee litigation was reported ongoing in mid-2026, and appellate review of the 2026 order remained possible as of August 2026.
California answered the appellate narrowing with a statute. SB 855 (Stats. 2020 ch. 151, approved 25 September 2020, effective 1 January 2021) requires commercial plans to make mental-health and substance-use medical-necessity determinations using the most recent criteria of the nonprofit professional association for the relevant clinical specialty — naming ASAM, the LOCUS and CALOCUS assessment scales, CASII and ECSII — and forbids applying different, additional, conflicting or more restrictive utilization review criteria. Where the Ninth Circuit held that ERISA plan terms do not themselves carry the standards requirement, the legislature imposed it categorically.
The sociotechnical reading
Read the whole atlas and you find the same architecture over and over: an instrument produces an output, an operator has too little time and too little standing to disagree with it, and the correction channels are too thin or too late. This case has that architecture too. What makes it worth its own file is that the instrument is a document, and the trial reached the room where the document was written.
That relocation changes what every control on the diagram is for. In a deployment built around a scored prediction, the governance question is usually whether anyone can see the score's error rate and act on it. Here there is no error rate to see, because there is nothing to be wrong in the way a prediction is wrong. The criteria said what they said, every reviewer applied them, and the deviation from generally accepted standards was a property of the text. That is why this network's centre of gravity is a record store — the annual edition — and why its widest pathway runs from that store into the determination rather than from any case into it. It is also why the two width asymmetries on the board are findings rather than styling. The criteria reach the determination at full width while the member's own clinical information reaches it at the narrowest rung, and that gap is the eighth declared deviation stated as a picture: mandatory prerequisites in place of a multidimensional assessment of the individual. The reviewing clinician's thirty-minute call with the treating physician is the one clinical input from outside the company, and it is drawn thin because the court found the testimony that it could change the outcome not credible.
The oversight story here is unusually precise, and it is not the usual one. Most governance failures in this catalogue are absences: no committee, no audit, no cadence. This deployment had a committee. It met, it minuted, it could withhold approval, and it approved every edition. What the court found is not that the governance was missing but that it was constituted out of the very pressure it existed to check — the finance function was seated inside it, benefit-expense impact was modelled before adoption, and liberalizations waited on a green light. So the board draws the committee's reinforcing pathways wider than its own check: what it approves reaches the desks, its sign-off reaches backwards into what the drafters even propose, and its gate on revisions sits at the floor. An oversight body whose composition embeds the pressure it should resist is a governance shape, not a governance gap, and the diagram would lie if it drew the seat empty.
Around that sit two more checks that were present and defeated, and the difference between them matters. The state criteria mandates were real, enforceable law: four states required external professional-society criteria, and the court eventually enforced all four. The mechanism that defeated them during the class period was not the mandate's weakness but the answer given to it — the crosswalks telling Connecticut regulators that all three ASAM residential levels were included, which the operative judgment declares the company knew to be false. A conformance check reads what it is told. Meanwhile the member's internal appeal existed, and the trial court found pursuing it to exhaustion would have been futile, because the appeal applied the same criteria that produced the denial. A second read against the same rulebook is not a second read.
Then there is the correction that landed, and its shape is the reason this case belongs at the end of a chapter rather than the start of one. The channel that finally fired was a federal court, and it took twelve years, a ten-day trial, three appellate dispositions, a granted mandamus and a remand to get there. What arrived is aimed at authorship: a declaratory judgment that the breach was willful and systematic, a permanent bar on the tainted editions, and five years of court-supervised accuracy with jurisdiction retained. What did not arrive is on the same page and has to be carried with equal weight. The class-wide wrongful-denial theory failed. The Ninth Circuit held that neither ERISA nor the plans required the guidelines to be coextensive with generally accepted standards, and that reprocessing was not appropriate equitable relief; a mandamus then compelled judgment for the company on the benefits claim. The reviewing layer disciplined the correcting layer, in as many words. So the correction runs one way only: corrected criteria forward, uncorrected outcomes backward. Roughly 67,000 coverage determinations stand.
That asymmetry is the governance lesson, and it is why the network's own numbers put the litigation channel's per-item power above its reach. A correction that is strong and rare and late is not the same instrument as one that is weak and frequent and early, and averaging them into a single "oversight exists" would erase exactly what this record teaches. The outermost loop makes the same point from the other side: where the appellate holding said plan terms could not carry the standards requirement, California wrote a statute that carries it categorically, and named the professional associations outright. The fastest correction in this whole record was legislation — which is a strange sentence to write, and it is what the twelve years mean.
Two boundaries hold, as they must. Served members are not modelled: no coverage decision, level of care, clinical judgment, health outcome or financial outcome for any person is computed from anything on this diagram, and the class-scale figures are recorded litigation facts — roughly 67,000 coverage determinations for roughly 50,000 people, because a member can have more than one denial, with the characterization that about half were children or adolescents carried as plaintiff-side reporting. And the structural inequity this case is about runs between benefit categories, not between subpopulations of served people: behavioral-health coverage governed by criteria a court found more restrictive than generally accepted standards is a parity finding, and no disaggregated measurement of denial or overturn rates by any characteristic of a served person exists in this record. The diagram does not manufacture the second from the first.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.