PAN Lab example
Cost-Proxy Care Stratification
The score that was right about the wrong thing: a cost-trained care model
A commercial model scores every patient in a health system's risk contracts from last year's claims, and the top slice of the ranking is offered a scarce care-management program. Modeled on the cost-proxy stratification case an independent team dissected in 2019. The model is accurate. That is the problem. Its training label is next year's cost, its stated purpose is finding the sickest patients, and the two come apart by race: less is spent on Black patients at the same level of illness, so at the same score Black patients carried 4.8 active chronic conditions against 3.8. Every routine accuracy check is computed against cost, so every routine check passes. Watch three things: the store the model reads (spending history, standing in for need), the store it never reads (the clinical record, where the severity signal lives), and the one experiment in the record — changing only the training label cut one bias measure by 84 percent, in a predictor that never reached production. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. With every tool the Lab currently offers, no affordable combination brings this system inside the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Cost-proxy-class care-management stratification score network: 10 components and 24 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 11 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the cost-proxy stratification pattern documented in the cost-proxy-care-stratification case file — not a reconstruction of the actual product. Its defining feature is target misspecification rather than model error: the production model is accurate and well calibrated across race on the target it was trained on, next year's cost, and the target is the wrong proxy for the deployment's stated purpose, so every accuracy check computed on that target passes by construction.
- baseline
Two model nodes are drawn because the record documents two predictors: the production cost-label score, and a relabelled health-and-cost index the evaluators and the manufacturer built jointly on the same sample, the same features with race still excluded, and the same training process, changing only the label. The relabelled predictor was experimental and holdout at publication and made no enrollment decisions, which is why its write path into the register is drawn at zero — the honest topology of a remedy that exists and is not in production.
- baseline
The clinical record's read path into the production model is drawn at zero because the record documents its absence, not because the shape looked bare: the production model's features are prior-year claims — demographics, insurance type, diagnosis and procedure codes, medications and detailed costs — and it does not consume the laboratory values or vital signs on which the disparity at equal score was demonstrated. The store that could correct the label is the store the model has no read path into, and this network draws that as topology rather than asserting it as prose.
- baseline
The claims store's edge into the production model is the widest in this network because one store is simultaneously the feature source and the label source, so a disparity in what was historically spent re-enters the predictor on both sides at once. The loop closes through the operators: program activity is billed by the care team, an enrollment becomes a billing fact, and the billing record is next year's feature set. The study ran experiments bounding the program's effect on the health measures used, so the replication edge is drawn thin rather than wide.
- baseline
Demand reads high and capacity very low from the documented shape of the program, not from a default. The program is documented as expensive and scarce — dedicated nurses and extra appointment slots, which is why a targeting model is used at all — and enrollment covered 1.3 percent of observations while everyone above the 55th percentile was surfaced to a physician. Capacity reads very low because no human process stratifies a 50,000-patient panel each period, and the one designed human channel was measured: it moved the enrolled cohort's composition by about one percentage point against race-blind sampling, where a different label would have moved it by about eight.
- baseline
The two threshold channels are drawn as reads of the register, because that is where the record puts them: the score is recorded, the two percentiles are applied to it, and the register is the record the care team and the referring physicians work from. Their mechanisms differ and both are documented: the referral screen covers everyone above the 55th percentile and is bounded by the physician's discretion to consider, while the auto-identification covers the top 3 percent and reaches the care team as a worklist rather than a question. The team's read is the strongest store read in the network because the auto-identified list is its entire working frame — and a patient the score never surfaced appears nowhere in it.
- baseline
The production model's self-check is drawn present at the low rung, and this is deliberate: a generation-time accuracy check exists and is routinely performed — the industry evaluation of the ten most widely used algorithms of this class used cost prediction as its accuracy metric — and the model passes it, because realized costs at equal predicted risk were roughly equal across race. The check is real, running, and aimed at the quantity the label already satisfies. What it is pointed at, not whether it runs, is this deployment's finding.
- baseline
The relabelled model's cross-check on the production score is drawn at the low rung because it happened once, as an experiment: setting a differently-labelled predictor beside the production score on the same sample cut excess active chronic conditions in Black patients conditional on score from 48,772 to 7,758. Those are counts of excess conditions — a bias measure, never an error rate — and the 84 percent figure is a reduction in that measure, on the manufacturer's national dataset of 3,695,943 commercially insured patients.
- baseline
The reviewer is wired in through its own inbound pathways, per the corpus idiom, and each is drawn from its documented reach: the production model's predictions reached the evaluators as one health system's dataset for 2013 to 2015 (the whole of their access, so the low rung), the relabelled predictor reached them as a joint holdout evaluation, and the clinical record reached them as the raw material of a constructed ground truth. Their outbound check on the deploying side is drawn present at the low rung because the disclosure demonstrably carried — replication, confirmation, a joint rebuild — through one study, with no standing cadence and no authority to stop anything.
- baseline
The enrollment worklist is drawn as an edgeless mediator per the corpus idiom: the 97th-percentile threshold produces an auto-identified queue each enrollment period, and the care team works it down. No enforcement component is drawn — the downstream action in this record is a care offer worked by a group of staff, not a penalty system replicating the record — and no external input-source node is drawn, because the claims history is inside the documented feedback loop rather than an external feed.
- baseline
The regulators are not drawn as a node. The New York DFS and DOH joint letter of October 25, 2019 was addressed to the vendor's parent company and demanded that it investigate and demonstrate the algorithm is not racially discriminatory or cease using it; 45 CFR 92.210 places a standing identify-and-mitigate duty on covered entities' use of patient care decision support tools. Both act on whether the tool may be used and what must be measured about it, from outside the deployment boundary — they are carried in the scenario's levers and the case file, not as a decorative node.
- baseline
Naming follows the attribution in the record. The Science study anonymized the product and manufacturer. The same day it published, New York's financial-services and health regulators wrote to UnitedHealth Group naming 'Optum's data analytics program, Impact Pro', and press reports carried the same identification with the company's response. This network's copy reports that attribution and never asserts the identification as a finding of the study itself. No publicly documented resolution of the New York inquiry had been located as of 2026-08.
- assumed
Granularity. This board draws the deployment at the coarsest granularity at which every documented mechanism is still separately visible. Where the record documents one act, the diagram draws one pathway: the referral packet and the auto-identified list are reads of the register rather than second copies of the model output, the analytics team's reading of the score distribution and its setting of the two percentiles are one configuration act at the register, the physician's handoff to the care team is the enrollment decision the register records, and the two chart pathways are each drawn in the direction the record makes consequential with the other half narrated alongside it. Every figure, mechanism and documented absence that a folded pathway carried is stated on the pathway that absorbed it; nothing was dropped to make the diagram smaller.
- assumed
Served patients are not in the dynamics. No patient outcome, enrollment decision, illness burden, cost or mortality figure is computed from anything drawn here. The study's counterfactual — that remedying the disparity would raise the fraction of Black patients receiving additional help from 17.7 to 46.5 percent — is a statement about the composition of the selected cohort, never a selection rate on a subpopulation, and it lives in the case file as a recorded external figure, not on this diagram.
What this example does not show
- No patient outcome is modeled. The Lab reads institutional propagation only; the patients being scored are boundary-only, and the chronic-condition gap, the cost calibration, the enrollment composition figures and the severity differences live in the case file as recorded external observations from one published evaluation, never computed on this diagram.
- The study's counterfactual — 17.7 percent to 46.5 percent — describes the composition of the group receiving additional help under a remedied ranking. It is the share of that selected cohort who are Black patients, never a rate of selection for any subpopulation, and this scenario does not present it as one. Likewise 48,772 and 7,758 are counts of excess active chronic conditions, a bias measure rather than an error rate, and 84 percent is a reduction in that measure.
- The Science study anonymized the product and the manufacturer. The identification of the product rests on the New York DFS and DOH joint letter of October 25, 2019 — which named 'Optum's data analytics program, Impact Pro' in a demand to UnitedHealth Group — and on contemporaneous press reports carrying the same identification with the company's response. This scenario reports that attribution and does not assert the identification as a finding of the study. No publicly documented resolution of the New York inquiry had been located as of August 2026.
- The relabelled predictor was an experimental, holdout model at publication, with no production write path; the remedy in this record was built and measured, never shipped, and no independent evaluation of it by anyone other than its authors has been located. The measured figures come from one large academic health system's cohorts, 2013 to 2015; the roughly 200 million figure is an industry estimate for the product class, not a measurement of this product.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A commercial population-health risk score used to target a scarce high-risk care management program at a large academic hospital (studied 2013-2015) was trained to predict total medical expenditure in the following year while the deployment's stated purpose was to identify the patients with the greatest health needs, with race excluded from its features. The independent evaluation found the model well calibrated for cost across race — roughly equal realized following-year costs at every level of predicted risk, 5,147 versus 4,995 dollars at the median score — while at the 97th-percentile auto-identification threshold Black patients carried 26.3 percent more active chronic conditions than White patients at the same score (4.8 versus 3.8, P < 0.001). Patients above the 97th percentile were automatically identified for enrollment and those above the 55th were referred to their primary care physician; enrollment covered 1.3 percent of observations, and commercial tools of this class are applied to roughly 200 million people in the US each year per industry estimates cited in the study.
empirical- Academic Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464): 447-453, DOI 10.1126/science.aax2342 (open deposit, UC Berkeley Previously Published Works; Europe PMC record MED/31649194) https://escholarship.org/uc/item/6h92v832
- Peer-reviewed Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). "Dissecting racial bias in an algorithm used to manage the health of populations." Science 366(6464): 447-453. DOI 10.1126/science.aax2342. https://europepmc.org/article/MED/31649194
The deployment's designed human check — physician referral above the 55th percentile — was measured: realized program enrollment was 19.2 percent Black against 11.9 percent Black in the whole sample, while simulated race-blind sampling within bins of the same score would have produced 18.3 percent (P = 0.8348 against the observed figure) and sampling on predicted chronic-condition count within bins would have produced 26.9 percent. The study's abstract states that remedying the disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5 percent — a figure describing the composition of the helped cohort under a counterfactual ranking, not a rate of selection for any subpopulation.
empirical- Academic Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464): 447-453, DOI 10.1126/science.aax2342 (open deposit, UC Berkeley Previously Published Works; Europe PMC record MED/31649194) https://escholarship.org/uc/item/6h92v832
- Peer-reviewed Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). "Dissecting racial bias in an algorithm used to manage the health of populations." Science 366(6464): 447-453. DOI 10.1126/science.aax2342. https://europepmc.org/article/MED/31649194
Per the Science study, the manufacturer independently replicated the analysis on its national dataset of 3,695,943 commercially insured patients, confirmed the finding, and jointly rebuilt the predictor with the study's authors using the same sample, the same features with race still excluded, and the same training process, changing only the training label to an index combining health prediction with cost prediction. One measure of predictive bias — excess active chronic conditions in Black patients conditional on risk score — fell from 48,772 to 7,758, an 84 percent reduction in that bias measure (counts of excess conditions, not error rates). The relabelled predictor was experimental and holdout at publication, with no production role in enrollment decisions, and no independent evaluation of it by anyone other than its authors has been located. The study's authors later generalized the remediation practice in the Algorithmic Bias Playbook (Chicago Booth Center for Applied AI, 2021), which centers label-choice bias and, like the study, does not name the product.
empirical- Academic Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464): 447-453, DOI 10.1126/science.aax2342 (open deposit, UC Berkeley Previously Published Works; Europe PMC record MED/31649194) https://escholarship.org/uc/item/6h92v832
- Reference Obermeyer, Z., Nissan, R., Stern, M., Eaneff, S., Bembeneck, E. J., & Mullainathan, S. (2021, June). Algorithmic Bias Playbook. Center for Applied AI at Chicago Booth https://www.chicagobooth.edu/-/media/project/chicago-booth/centers/caai/docs/algorithmic-bias-playbook-june-2021.pdf
The Science study anonymized the product and manufacturer. On its publication day, October 25, 2019, the Superintendent of the New York State Department of Financial Services and the Commissioner of the New York State Department of Health jointly wrote to UnitedHealth Group's chief executive, naming 'Optum's data analytics program, Impact Pro', footnoting the study, and demanding that the company immediately investigate and demonstrate the algorithm is not racially discriminatory or cease using Impact Pro or any similar program, asserting insurers' responsibility under New York law regardless of who built the algorithm; contemporaneous press reports carried the same identification and the company's statement that the cost model 'was highly predictive of cost, which is what it was designed to do.' No publicly documented resolution of the New York inquiry has been located as of August 2026. Separately, 45 CFR 92.210 (adopted in the 2024 HHS section 1557 final rule) prohibits discrimination through the use of patient care decision support tools and imposes ongoing duties to identify tools employing inputs measuring race, color, national origin, sex, age, or disability and to mitigate the resulting risk — a standing duty that postdates the studied deployment and is carried for the duty it imposes, not as an adjudication of this case.
empirical- Government New York State Department of Financial Services and New York State Department of Health, joint letter to UnitedHealth Group Incorporated, 25 October 2019 https://www.dfs.ny.gov/reports-and-publications/comment-letters/dfs-doh-joint-letter-uhgi-20191025
- Government 45 CFR 92.210, Nondiscrimination in the use of patient care decision support tools (U.S. Department of Health and Human Services, section 1557 implementing regulation) https://www.ecfr.gov/current/title-45/section-92.210
- news Snowbeck, C. (2019, October 29). NY Regulators Probe for Racial Bias in Health-Care Algorithm. Star Tribune (via Government Technology syndication) https://www.govtech.com/health/NY-Regulators-Probe-for-Racial-Bias-in-Health-Care-Algorithm.html
- news Ledford, H. (2019). Millions of black people affected by racial bias in health-care algorithms. Nature (news), 574, 608-609 https://www.nature.com/articles/d41586-019-03228-6
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
All of them in context on the Clinical decision support & deterioration alerting domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Upgrade model — Improve the model
- Check with a second model — Cross-model verification
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Gate vendor updates — Vendor quality gate
- Mark AI-written records — Provenance labeling
- Gate record entries — Human-in-the-loop write gating
- Check copied records — Reconcile copied records
Documented case histories
- Cost-Proxy Care Stratification
- TREWS sepsis early-warning system
- Advance Alert Monitor (AAM) deterioration model
- Sepsis Watch deep-learning detection system
- Proprietary EHR sepsis model (external validation)
- nH Predict Utilization Review
- CA-CDS Child Abuse Alerting
- IDx-DR Autonomous Screening
- Viz.ai LVO Stroke Triage
- IBM Watson for Oncology
- OPTN eGFR Waiting-Time Correction
- Practice Fusion Pain CDS
- UBH Level of Care Guidelines (Wit v. UBH)
- EviCore by Evernorth: the review threshold
- Cigna PxDx