Domain Atlas / Clinical decision support & deterioration alerting
Cost-Proxy Care Stratification
A commercial population-health risk score used to target a scarce high-risk care management program at a large academic hospital (studied 2013-2015) was trained to predict total medical expenditure in the following year while the deployment's stated purpose was to identify the patients with the greatest health needs, with race excluded from its features. The independent evaluation found the model well calibrated for cost across race — roughly equal realized following-year costs at every level of predicted risk, 5,147 versus 4,995 dollars at the median score — while at the 97th-percentile auto-identification threshold Black patients carried 26.3 percent more active chronic conditions than White patients at the same score (4.8 versus 3.8, P < 0.001). Patients above the 97th percentile were automatically identified for enrollment and those above the 55th were referred to their primary care physician; enrollment covered 1.3 percent of observations, and commercial tools of this class are applied to roughly 200 million people in the US each year per industry estimates cited in the study.[2]
What happened
Between 2013 and 2015, a large academic hospital ran its high-risk care management program — dedicated nurses, extra primary care appointment slots, and other deliberately scarce resources — on a commercial population-health risk score. Every primary care patient in risk-based contracts was scored each enrollment period from the prior year's insurance claims: demographics, insurance type, diagnosis and procedure codes, medications, and detailed costs, with race explicitly excluded from the features. Patients above the 97th percentile of the score were automatically identified for enrollment; patients above the 55th percentile were referred to their primary care physician, who received contextual data and was asked to consider whether the patient would benefit. Program enrollment covered 1.3 percent of observations. In the framework the National Academies set out for integrating social care into health care, this is an assistance activity at its high-touch end — intensive case management — and the score decided who reached it.
In October 2019, Obermeyer, Powers, Vogeli, and Mullainathan published in Science the dissection of that deployment, with access to the algorithm's inputs, outputs, and objective function — a rare vantage. The mechanism they isolated was the training label. The model was trained to predict total medical expenditure in the following year; the deployment's stated purpose was to identify the patients with the greatest health needs; and those two quantities come apart by race, because unequal access to care means less is spent on Black patients at the same level of illness. The evaluation measured both halves on the same cohorts — 6,079 self-identified Black patients (11,929 patient-years) and 43,539 self-identified White patients (88,080 patient-years). On cost, the model was well calibrated across race — calibration meaning the predicted figure really does match what actually followed: roughly equal realized following-year costs at every level of predicted risk ($5,147 versus $4,995 at the median score; $35,541 versus $34,059 in the top 5 percent). On health, at the same score, Black patients were substantially sicker on every marker examined: at the 97th-percentile auto-identification threshold they carried 26.3 percent more active chronic conditions (4.8 versus 3.8, P < 0.001), with severity gaps including 5.7 mmHg of systolic pressure and 0.6 percentage points of glycated hemoglobin. The abstract's counterfactual: remedying the disparity would raise the percentage of Black patients receiving additional help from 17.7 to 46.5 percent — a statement about the composition of the helped cohort, not a selection rate. The study also measured the deployment's designed human check. Realized enrollment was 19.2 percent Black against 11.9 percent in the whole sample; simulated race-blind sampling within score bins would have produced 18.3 percent (P = 0.8348 against the observed figure); sampling on predicted chronic conditions would have produced 26.9 percent. The physician screen redressed a small part of the bias, and far less than a different label would.
The aftermath moved on three tracks at once. The Science study anonymized the product and its manufacturer — but on the day it published, October 25, 2019, the Superintendent of the New York State Department of Financial Services and the Commissioner of the New York State Department of Health jointly wrote to UnitedHealth Group's chief executive, naming "Optum's data analytics program, Impact Pro," footnoting the study, and demanding the company immediately investigate and demonstrate that the algorithm is not racially discriminatory or cease using Impact Pro or any similar program — asserting that insurers are responsible under New York law whether or not they built the algorithm. Contemporaneous reporting (Star Tribune, October 29) carried the same identification and the company's statement that "the cost model within Impact Pro was highly predictive of cost, which is what it was designed to do." No publicly documented resolution of the New York inquiry has been located as of August 2026. On the second track, the manufacturer independently replicated the analysis on its national dataset of 3,695,943 commercially insured patients, confirmed the finding, and jointly rebuilt the predictor with the study's authors — same sample, same features with race still excluded, same training process, changing only the label to an index combining health prediction with cost prediction. One measure of predictive bias, excess active chronic conditions in Black patients conditional on score, fell from 48,772 to 7,758: an 84 percent reduction, in a predictor that was experimental and holdout at publication and was not the tool making enrollment decisions. On the third track, the finding generalized: the authors' Algorithmic Bias Playbook (Chicago Booth Center for Applied AI, 2021) distilled a label-choice-centered audit practice from applied work with dozens of organizations, and in 2024 the HHS section 1557 rule created 45 CFR 92.210, a standing federal duty on covered entities to identify patient care decision support tools that pose discrimination risk and make reasonable efforts to mitigate it. The rule postdates the studied deployment and adjudicates nothing about it; what it changes is who now has to ask the question this study asked once.
The sociotechnical reading
Every other clinical-decision-support case in this atlas is an event predictor that can be wrong about its event: a sepsis model can miss sepsis, and an accuracy metric computed on its own target can in principle catch it. This case is a different failure class, and the atlas carries it for exactly that reason: the model was right, on the target it was given, and the target was the wrong one. That inversion breaks the standard governance toolkit at the joint everyone leans on. The industry's routine evaluation of this product class used cost prediction as its accuracy metric — so the check applied at generation was, by construction, the check the model passes. The deployment's own monitoring watched utilization and cost, the quantity the label already satisfies. The operators checking a score against the deployment's claims-derived record got back "correct" precisely when the target was wrong, because the claims history is an accurate account of what was spent and a distorted measure of what was needed — the missing spending of a patient whose access was constrained leaves no record to audit. And the deployment's designed human override, the physician referral screen, was measured and moved the enrolled cohort's composition by about one percentage point against race-blind sampling, where a different label would have moved it by about eight. Checking more often against a record that measures spending rather than need is arithmetic that does not compound. The correction that worked came from outside the deployment's record system entirely: an external team constructed an independent ground truth by linking predictions to diagnoses, laboratory values, and vital signs — the store the production model never read — and the one lever that moved the bias measure by 84 percent was the training label, changed with everything else held constant.
The map draws all of this as structure. The claims store feeds the production model on both sides at once — features and label — and the loop closes through the operators, because program activity is billed and billing is next year's training data; the clinical record, where the severity signal lives, has no read path into the production model; and the relabelled predictor, the one channel wired to the clinical record, has no write path into the register that enrollment runs on. The honest complications carry with the praise. The disclosure loop in this record worked unusually well — replication at 3.7 million patients instead of denial, a joint rebuild instead of litigation — and it still ended, on the public record, with an experimental holdout model, a product line the vendor's own materials still list, and no located resolution of the regulators' inquiry. The 17.7-to-46.5 counterfactual is a composition figure for the helped cohort and must never be read as anyone's chance of selection; 48,772 and 7,758 are counts of excess conditions, a bias measure and not an error rate; and the naming chain is itself part of the record's structure: the study anonymized the product, the New York regulators named it, the press corroborated the naming, and this atlas reports that attribution without asserting it beyond its sources. What the record does not establish also matters. The measured figures belong to one health system's cohorts across two years; the 200-million figure is an industry estimate for the product class; and no patient outcome anywhere in this case is computed by the Lab — the cohort statistics are recorded external observations from one published, peer-reviewed evaluation, with their own denominators, and the diagram propagates institutions, not patients.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library.