Skip to content

PAN Lab example

Epic Sepsis Model

Switched on before anyone checked: a proprietary sepsis model at scale

A proprietary model shipped inside the electronic health record (EHR) is switched on across hundreds of hospitals - before anyone independent checks whether it works. When they do, it catches a third of sepsis at about 109 alerts per true case. Modeled on the external-validation record of a widely deployed sepsis model. This one starts broken: the alert flood has hollowed out the clinician's attention, and no validation ever ran. Your job is the governance that should have come first.

Stylized model of a documented deploymentClinical decision support & deterioration alerting

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Proprietary-EHR-sepsis-class model at scale network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 7 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This org and the other sepsis early-warning deployment share one topology by ruling: both records describe the same machine - a model scoring a continuous live patient-data stream inside the electronic health record (EHR), alerting the treating clinician, under an external-validation function - and neither documents a component the other lacks, so drawing a difference would be invention. The live stream is drawn as its own inbound source at full strength because it is the record's defining input, and in this deployment it is also where the documented leakage lived: antibiotic orders arrived as live inputs, so the score partly restated decisions already made. Uniqueness between the twins is carried entirely by dynamics, which is where the two records genuinely diverge.

  • assumed

    The record-to-model pathway is drawn at full strength because the external validation found an undisclosed feature leak: the model used antibiotic-order data, so it was partly reading the treatment decision it was meant to predict. That is the sharpest store-to-model contamination in this catalogue and it is documented, not inferred - a model learning from the clinician action downstream of its own alert.

  • assumed

    The model-to-model pathway is drawn at full strength: one proprietary model, built and updated by a single vendor, deployed across hundreds of hospitals. Nothing in this catalogue is more correlated by construction, and the same blind spot is the same blind spot everywhere it runs.

  • assumed

    The independent-validation pathway is drawn faint, not empty, and this is what distinguishes this org from every deployment whose failure is an absent audit. An external academic validation over 38,455 hospitalizations was performed and published, and it was damning - poor discrimination, a third of cases caught, and an alert burden of about a hundred and nine per true case. The model nonetheless stayed deployed at scale until its vendor chose to overhaul it. The latent pathway here is therefore not the check but the authority to act on one: a finding that reaches nobody who can stop a deployment is not oversight.

  • baseline

    This models the deploy-at-scale-before-validation failure pattern documented in the Epic-class sepsis case file - not a reconstruction of the actual model. It is drawn strained at baseline by design (high alert-flood demand, fatigue-degraded correction) because that is the honest failure regime: a proprietary model switched on across hundreds of hospitals before independent validation, catching about a third of sepsis at roughly 109 alerts per true case.

  • assumed

    The alert flood is drawn on the model-to-staff pathway at full strength and the correction faint: at ~109 alerts per true case the overwhelming majority of firings are false, so the clinician's evaluation collapses into reflexive dismissal - the alert-fatigue mechanism by which a nominally-helpful tool becomes noise the workflow routes around. Aggressive tiering and alert-silencing are the honest correction, which the retuned model's own authors now recommend.

  • assumed

    The defining absence is drawn on the independent model check, empty at baseline: no independent validation ran before the model was deployed at scale, and it was shielded behind a vendor firewall from the scrutiny that would have surfaced its performance. The 2021 external validation and the 2026 retuned-model validation are that check firing years late; the retuned model's authors still urge local validation before trusting it - drawn here as the dormant oversight check.

  • assumed

    No patient or sepsis outcome is modeled here. This Lab reads institutional propagation only, and the patients being scored are boundary-only. The accuracy numbers, the ~109-alerts-per-case ratio, the undisclosed-features finding, and the retune live in the case file, and are never computed from anything in this diagram.

What this example does not show

  • No patient or sepsis outcome is modeled. The Lab reads institutional propagation only; the patients being scored are boundary-only, and the accuracy numbers, the ~109-alerts-per-case ratio, and the retune live in the case file, never computed on this diagram.
  • The strained baseline is a modeling choice representing the documented failure regime (deploy-before-validation, alert fatigue, vendor opacity); the specific alert ratio and accuracy figures are the external-validation record, not values derived from the diagram.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and switched on across hundreds of hospitals was externally validated in 2021 across 38,455 hospitalizations at an academic health system: it achieved an area under the curve of 0.63, identified only 33 percent of sepsis cases, and had a positive predictive value of about 12 percent, generating roughly 109 alerts for every true sepsis case — a real-world performance the vendor had not fully examined before selling the model, and which an investigation attributed in part to undisclosed features such as antibiotic-order data that inflated internal validation.

    empirical
    • Academic Wong, A., Otles, E., et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307
    • Investigative STAT News (2021, July 26). Epic's AI algorithms, shielded from scrutiny by a corporate firewall, are delivering inaccurate information on seriously ill patients. https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/
  • After external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, and reducing its reliance on antibiotic-order features. A 2026 multicenter prospective validation of the updated model across 227,091 encounters reported an area under the curve of 0.82 to 0.92 with positive predictive value of 0.13 to 0.26 and substantial between-site variability, and its authors urged local validation and alert-silencing strategies rather than trusting the model out of the box — a correction that arrived only after independent scrutiny of a model that had already been deployed at scale behind a corporate firewall shielding it from outside inspection.

    empirical
    • Investigative STAT News (2022, Oct 3). Epic overhauls popular sepsis algorithm criticized for faulty alarms. https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/
    • Peer-reviewed Wong, A., Currey, D., Schwinne, M., et al. (2026). Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model. JAMA Network Open. https://doi.org/10.1001/jamanetworkopen.2026.0181 https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2845595

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical decision support & deterioration alerting domain page.

Levers available here and the patterns behind them

Documented case histories