Clinical decision support & deterioration alerting
Machine-learning early-warning models that flag hospitalized patients for sepsis or clinical deterioration — the domain where an AI's measured benefit is real but runs entirely through the human loop it interrupts. The same alert that saves a life when a clinician confirms it in time becomes a source of fatigue when it fires a hundred times per true case; what separates the two is whether the confirmation workflow is resourced, whether the model was validated independently of the vendor who sells it, and whether anyone reconciles the alerts against the outcomes they were meant to change. The Lab networks here model only the deploying hospital — its models, clinicians, and records; the patients being scored sit outside the dynamics, and no clinical outcome is ever computed on a diagram.
Use cases
What AI is doing here
Sepsis & deterioration early-warning alerting
PredictiveMachine-learning models that continuously score hospitalized patients' electronic-health-record data and raise an alert when sepsis or clinical deterioration is predicted, for a clinician to evaluate and confirm — a signal whose measured benefit is contingent on a resourced human confirmation step.
Rapid-response & virtual-nurse escalation
PredictiveDeterioration alerts routed through a mediating escalation tier — a regional virtual-nurse desk or a rapid-response-team nurse — who screens the score and mobilizes bedside care, embedding the model inside a staffed workflow whose hidden coordination labor is what makes the alert actionable.
Proprietary EHR-embedded risk scores
PredictiveVendor-built risk models shipped inside a widely used electronic-health-record platform and switched on across many hospitals at once, where the model's real-world accuracy and alert burden may not be independently validated before deployment and the vendor's internal evaluation is shielded from outside scrutiny.
Case files
What has gone wrong and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
TREWS sepsis early-warning system
United States (Johns Hopkins Medicine; five hospitals across Maryland and Washington, DC)A machine-learning sepsis alert deployed across five hospitals of an academic health system, evaluated on 590,736 patients — the largest prospective study of its kind. Its mortality benefit was real but conditional: it accrued only to patients whose alert a provider confirmed within three hours, and the alert alone did nothing. The evaluation was prospective and peer-reviewed, but observational and run by the developer who commercialized the model.
Explore this deployment in the PAN Lab →Advance Alert Monitor (AAM) deterioration model
United States (Kaiser Permanente Northern California; 21 hospitals)An in-hospital deterioration model running around the clock across 21 hospitals, firing about twelve hours ahead of predicted deterioration. Its NEJM-measured mortality benefit is inseparable from where the alert goes: not to the bedside, but to a dedicated regional tier of critical-care virtual nurse consultants who screen every alert before escalating. The governance here is not a checkbox — it is a staffed, 24/7 subsystem with a payroll.
Explore this deployment in the PAN Lab →Sepsis Watch deep-learning detection system
United States (academic hospital; registered clinical trial NCT03655626)A deep-learning sepsis detector scoring every emergency-department patient every five minutes, fronted by rapid-response nurses on treatment-bundle timers. Its fault line is an authority split: the nurse who receives the alert is not the physician empowered to act on it. An independent ethnography found the system worked only because nurses did hidden, undervalued repair work to make a risk score actionable across a professional hierarchy.
Explore this deployment in the PAN Lab →Proprietary EHR sepsis model (external validation)
United States (proprietary model in a widely used EHR; external validation at an academic health system)A proprietary sepsis-prediction model shipped inside a common EHR and switched on across hundreds of hospitals — then externally validated to catch only a third of sepsis cases at roughly 109 alerts per true case, a real-world performance the vendor had not fully examined before selling it, shielded behind a firewall from outside scrutiny. This is the domain's failure arc: deployment at scale ahead of independent validation, alert fatigue, and vendor opacity — corrected only after outside criticism forced a retune.
Explore this deployment in the PAN Lab →nH Predict Utilization Review
United States — federal. Private Medicare Advantage plans administering a federal benefit under contract with CMS; the governance surface is 42 CFR 422.101(c) as amended by CMS-4201-F (effective January 1, 2024), the CMS HPMS FAQ memo of February 6, 2024, HHS OIG evaluation OEI-09-18-00260, the Senate Permanent Subcommittee on Investigations majority staff report of October 17, 2024, and federal class actions in the District of Minnesota (Lokken) and the Western District of Kentucky (Barrows)The payer side of clinical decision support: inside the largest Medicare Advantage insurer's post-acute coverage review, a vendor coordinator completes the nH Predict similar-patient stay estimator while the patient is still hospitalized, and subpoenaed committee minutes show the same testing that cut six to ten minutes from each review arriving together with a rise in the adverse-determination rate - and being approved anyway. The one correction with a measured effect, the appeal, overturned 83.2% of the denials that reached it in 2022; 9.9% of denials did.
Explore this deployment in the PAN Lab →Cost-Proxy Care Stratification
United States — one large academic hospital's high-risk care management program (studied 2013-2015); product class national (roughly 200 million people scored per year by commercial tools of this kind, per industry estimates cited in the study); regulatory events in New York State and, since 2024, under 45 CFR 92.210A commercial risk score used to ration a scarce care-management program was trained to predict next year's medical cost while its stated purpose was to find the sickest patients — and because unequal access means less is spent on Black patients at equal illness, the model was accurate and racially well calibrated on cost while Black patients at the same score carried 26.3% more chronic conditions. The independent dissection (Science, 2019), the manufacturer's own replication on 3.7 million patients, a same-day New York regulator demand letter, and an 84% reduction in one bias measure from changing only the training label — in an experimental, holdout predictor that was not the tool making enrollment decisions.
Explore this deployment in the PAN Lab →CA-CDS Child Abuse Alerting
United States — Pennsylvania and New York (UPMC network), Wisconsin (UW Health), New York (Northwell Health); receiving side governed by state statute in each jurisdiction under the federal Child Abuse Prevention and Treatment Act (CAPTA)A rule-and-text child-abuse alerting system built by the CA-CDS Consortium at UPMC Children's Hospital of Pittsburgh and disseminated onto two other health systems' commercial record platforms roughly doubled triggering at both and measurably raised mandated reporting — but its working end product crosses a boundary its builders cannot see across: a report an individual clinician may not withhold, landing in a state child protection agency's durable record that must respond, screens out more than it screens in nationally, and by statute owes nothing back to the chart the report came from.
Explore this deployment in the PAN Lab →Cigna PxDx
United States — a national commercial insurer administering roughly 18 million lives per the investigative record. Governance surfaces: ERISA (29 U.S.C. § 1132) in Kisting-Leung v. Cigna Corp. (E.D. Cal., No. 2:23-cv-01477-DAD-CSK), coordinated with Snyder v. The Cigna Group (D. Conn., No. 3:23-cv-1451-OAW); California Health & Safety Code § 1367.01(e) via the surviving state unfair-competition claim; the California Department of Managed Health Care enforcement action of October 8, 2025 against Cigna HealthCare of California, Inc. ($500,000, agreed corrective actions); and the House Energy and Commerce Committee document inquiry of May 2023. Reported but unresolved: scrutiny by the U.S. Department of Labor and by the California, Washington, and Delaware insurance regulators.The domain's inversion case: a claim-review system that decides nothing about care, because the care has already happened. A deterministic code screen compares the procedure a doctor billed against an in-house list of diagnoses the insurer deems acceptable for it; matches are paid, mismatches queue to a company physician who signs the denial. ProPublica and The Capitol Forum, computing from internal company records, reported over 300,000 payment requests denied through this method in two months of 2022 at an average of 1.2 seconds each. Cigna disputes that characterization and has published no substitute figures. In October 2025 California's managed-care regulator fined the state plan entity $500,000, finding claims denied without physicians conducting clinical reviews first, under a review process that differed from the one it had on file.
Explore this deployment in the PAN Lab →EviCore by Evernorth: the review threshold
United States — national. EviCore by Evernorth (eviCore healthcare MSI, LLC d/b/a eviCore healthcare), a Tennessee-domiciled utilization-review entity licensed state by state (UR license 2552628) and owned by The Cigna Group since 2018, operating inside the Evernorth health-services arm. The governance surface is state utilization-review licensure rather than any federal regulator of the vendor: the Connecticut Insurance Department market conduct examination and stipulation and consent order, Docket MC 24-15 (February 5, 2024, allegations admitted, $16,000 fine, corrective-action report due in 90 days, violations of Conn. Gen. Stat. 38a-591b and 38a-591d and Reg. 38a-591-8); state-published denial data from Arkansas and Vermont Medicaid; a 2018 CMS audit that reached the vendor through its insurer client HCSC; and, as context rather than jurisdiction, HHS OIG evaluation OEI-09-18-00260 and the Senate Permanent Subcommittee on Investigations majority staff report of October 17, 2024, both of which examine insurers and not this vendorThe domain's routing case: an algorithm that denies nothing and still governs how much gets denied. The largest delegated prior-authorization vendor scores each request against criteria it writes itself; requests above an operating threshold are approved with no clinical review, and only the physicians below it may issue a denial. ProPublica and The Capitol Forum, working from internal documents and five former employees, reported in October 2024 that the threshold is adjustable and that insiders called it the dial. EviCore and Cigna dispute that characterization, saying the algorithms exist only to accelerate approval of appropriate care. The one enforcement loop that has closed is small and admitted: a Connecticut market-conduct examination of 196 files ended in a February 2024 consent order and a $16,000 fine for utilization-review compliance violations, and it read files rather than thresholds.
Explore this deployment in the PAN Lab →IBM Watson for Oncology
United States vendor (IBM Watson Health; Memorial Sloan Kettering Cancer Center as training partner, New York) with a global deployment footprint — roughly 50 hospitals on five continents by 2017, primary markets India, South Korea, China, Thailand, and Mongolia, plus a Danish pilot that declined adoption; the separate MD Anderson Oncology Expert Advisor project sits under the University of Texas System (special procurement review reported November 2016, public February 2017)IBM sold hospitals on five continents a cancer treatment advisor marketed as machine reading of the medical literature, whose recommendations were in fact computed from a knowledge base of synthetic cases written by a few specialists per cancer type at one New York hospital — a provenance the buyers were not told. Its own internal reviewers recorded examples of unsafe and incorrect recommendations in mid-2017 and global selling continued for years; no published study ever measured a patient outcome; and the arc closed with a private-equity divestiture rather than any safety verdict. Its separately owned twin project at MD Anderson consumed more than 62 million dollars and was benched by a state procurement audit that expressly declined to judge the medicine.
Explore this deployment in the PAN Lab →IDx-DR Autonomous Screening
United States — federal medical-device regulation (FDA De Novo DEN180001, creating device class 'retinal diagnostic software', 21 CFR 886.1100, product code PIB) and Medicare payment policy (AMA CPT code 92229; CMS CY2022 Physician Fee Schedule final rule); deployed in US primary care. Separate earlier EU-version validation in the Netherlands (Hoorn Diabetes Care System) and a later independent evaluation in Germany (Karlsburg Diabetes Hospital).In April 2018 the FDA authorized IDx-DR — since renamed LumineticsCore — the first diagnostic in any field of medicine whose clinical decision is rendered by software alone: a locked classifier that reads two retinal photographs taken by a clinic assistant who has never done ocular imaging, and returns one of two messages — refer to an eye care professional, or rescreen in twelve months — with no clinician interpreting the image or the result. Eight years on this is the atlas's benefit-forward anchor and its cleanest autonomy-boundary-by-design case: the human interpretive check was not eroded here, it was removed on purpose, in public, by a regulator, with the compensating controls named out loud. The measured benefit is real and replicated and the adverse-event docket is empty. The honest catch sits in the compensating control itself — in independent real-world use the image-quality gate that makes the design defensible declined a quarter of patients, disproportionately the older, cataract-prone people a screening programme most needs to reach.
Explore this deployment in the PAN Lab →OPTN eGFR Waiting-Time Correction
United States — the national organ allocation network. The HRSA-contracted Organ Procurement and Transplantation Network (OPTN), operated by UNOS, applied across all US kidney transplant programs (roughly 230 active programs). Policy record: the race-neutral eGFR requirement (OPTN Board unanimous 27 June 2022, effective 27 July 2022), the Waiting Time Modifications policy (Board unanimous 5 December 2022, effective 5 January 2023, attestation deadline 3 January 2024), and the Monitor Ongoing eGFR Modification Policy Requirements update (Board June 2025, effective 10 September 2025, program completion due 11 September 2026)For over a decade the standard equations for estimating kidney function multiplied the result upward for any patient recorded as Black — a coefficient of 1.159 in the 2009 CKD-EPI equation, against a measured median overestimate of 3.7 mL/min/1.73m2 — and a candidate needs an estimate of 20 mL/min or lower to begin accruing kidney waiting time. The US organ network did something this atlas rarely records: it prohibited the race-inclusive calculation in July 2022, then ordered every kidney program to recompute affected candidates' history without the coefficient and backdate the waiting time into the live allocation registry. By the one-year report, 14,701 modifications had been processed at a median of 1.7 years each and all 230 active kidney programs had attested. Then an independent national evaluation found the execution had varied significantly between transplant centers — and the governance response was to tighten, not to close the file.
Explore this deployment in the PAN Lab →Practice Fusion Pain CDS
United States — federal (U.S. Attorney's Office, District of Vermont; United States v. Practice Fusion, Inc., No. 2:20-cr-00011-wks), with a companion federal civil False Claims Act settlement (DOJ, HHS Office of Inspector General, Defense Health Agency), separate state Medicaid settlements, and a related guilty plea by the sponsor in the District of New JerseyA free cloud record platform sold the content of a point-of-care pain alert to an opioid manufacturer's marketing department for $959,700, let the sponsor's marketers propose edits to the trigger logic and the treatment-option list, and displayed the result more than approximately 230 million times over two and a half years — while its own analyses, delivered to the sponsor as the reporting it had bought and never to the prescribers being alerted, showed extended-release opioids to be the least effective of the listed options at lowering pain. The alert never named a drug. It did not malfunction. It is the atlas's clearest case of a system working exactly as designed, for a party the clinician could not see.
Explore this deployment in the PAN Lab →Viz.ai LVO Stroke Triage
United States — FDA De Novo DEN170073 (13 February 2018), creating the radiological computer-aided triage and notification class at 21 CFR 892.2080, product code QAS; CMS FY2021 hospital inpatient prospective payment final rule (85 FR 58432, 18 September 2020), new-technology add-on payment billed via ICD-10-PCS 4A03X5D, renewed for FY2022; deployed across US hospital stroke networksA detector reads every stroke-code CT angiogram as it leaves the scanner and, on a suspected large-vessel occlusion, pushes an alert with a compressed preview to the neurointerventional team's phones — while the radiologist's read continues beside it, unchanged and holding the diagnostic authority. The FDA authorization that created this device class in 2018 turns on exactly that parallelism, and CMS then scaled the deployment with an add-on payment of up to $1,040 a case. The system denies nothing and blocks nothing; its failure surface is alarm burden, a reading queue reordered at someone else's expense, distal occlusions the detector measurably misses, and an adoption curve that rose and fell with the payment rather than with need.
Explore this deployment in the PAN Lab →UBH Level of Care Guidelines (Wit v. UBH)
United States — federal. Wit v. United Behavioral Health, N.D. Cal. No. 14-cv-02346-JCS (related Alexander v. United Behavioral Health, No. 14-cv-05337): a ten-day ERISA bench trial before Chief Magistrate Judge Joseph C. Spero; Ninth Circuit Nos. 20-17363/20-17364 and 21-15193/21-15194 and mandamus No. 24-242. The criteria mandates of Connecticut, Illinois, Rhode Island, and Texas were adjudicated within the case; California SB 855 (2020) is the legislative echo.A behavioral-health claims administrator wrote the coverage criteria it applied to its own members, reissued them every year, and seated its Finance and Affordability representatives on the committees that approved them. After a ten-day bench trial a federal court found the 2011-2017 editions significantly and pervasively more restrictive than generally accepted standards of care, in eight enumerated ways, and found that the financial incentives had in fact infected the guideline development process. It is the atlas's best-documented case in which the rule itself was the defect — and its twelve-year appellate history is part of the record: the class-wide wrongful-denial theory failed, no coverage request was ever reprocessed, and what stands is a judgment about how the criteria were written.
Explore this deployment in the PAN Lab →System map
Who is in the system and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
- Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
Dominant pressures
- Workload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure. Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
Governance
Questions leaders should be asking
- 1. This model's measured benefit ran entirely through providers confirming its alerts in time — so is the confirmation step actually resourced, or is the benefit being claimed for a review the workload cannot sustain?
- 2. Who validated this model, and were they independent of the party selling it? A prospective, peer-reviewed evaluation run by the developer is still the strongest number produced by the most interested party.
- 3. How many alerts fire per true case, and who measures whether the clinicians being interrupted have started tuning the alarm out — the fatigue that turns a working tool into background noise?
- 4. Is anyone reconciling the alerts and the confirmations against the outcome the system exists to change, or only against how often it fired — and would a model that drifted out of calibration be caught before or after a bad patient outcome?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With Paramerge