Skip to content

PAN Lab example

TREWS sepsis early-warning system

The alert that works only when confirmed: a sepsis early-warning model

A machine-learning model scores every inpatient for sepsis and alerts a clinician to evaluate. Modeled on TREWS. Its measured mortality benefit was real - but it accrued only to patients whose alert a provider confirmed within three hours, and the alert alone did nothing. So watch two things the accuracy number cannot see: whether the confirmation step stays resourced and live, and who validated the model - because the strongest evidence here was produced by the party that built and sells it.

Stylized model of a documented deploymentClinical decision support & deterioration alerting

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the TREWS-class sepsis early-warning alert network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This org and the other sepsis early-warning deployment share one topology by ruling: both records describe the same machine - a model scoring a continuous live patient-data stream inside the electronic health record (EHR), alerting the treating clinician, under an external-validation function - and neither documents a component the other lacks, so drawing a difference would be invention. The live stream is drawn as its own inbound source at full strength because it is the record's defining input: the system monitored 590,736 patients, scoring as vitals and labs land, which is why an alert can precede the picture a person would have assembled from the chart. Uniqueness between the twins is carried entirely by dynamics - the confirm-within-three-hours engagement loop and the exercised prospective evaluation on this side, the alert flood and the vendor-opacity fight on the other.

  • assumed

    The benefit here is conditional on one pathway: the prospective multi-site study found the mortality reduction among sepsis patients whose alert a provider confirmed within three hours, so the provider-to-model confirmation pathway carries the benefit, not the model. It is drawn moderately rather than at full strength because the companion adoption paper documents that confirmation varies by provider experience, unit culture, and alert context - real, and unevenly so. Demand and capacity are drawn level because this deployment is not documented as overwhelmed; what varies here is engagement, not headcount, and that is the honest difference between this org and the flooded ones in its domain.

  • assumed

    It is drawn faint rather than empty: a prospective, multi-site, peer-reviewed evaluation across 590,736 monitored patients exists for this system, so drawing it as absent would misstate the record. But it is deliberately not drawn higher, because that evaluation is observational and developer-led - built at the deploying institution and commercialized through a company founded by its principal investigator. A published self-evaluation is a real read and not an arm's-length one, and the gap between those two is what the remaining headroom on this pathway represents.

  • assumed

    This models the confirmation-conditional sepsis-alert pattern documented in the TREWS case file - not a reconstruction of the actual model. Its defining feature is that the measured mortality benefit ran entirely through providers confirming alerts within three hours; the alert on its own did not reduce mortality.

  • baseline

    Unlike most tools in this Atlas, the human loop here is drawn active, not inactive: the provider's evaluate-and-confirm step carried the entire measured benefit. But a companion adoption study found that confirmation varied with provider experience, unit culture, and alert context, so the confirmation rate is a property of the workflow, not a fixed property of the tool - which is why the benefit is losable without changing the model, by letting the confirmation step erode.

  • baseline

    This shape's defining absence is independence: the independent model check is drawn empty at baseline because the evaluation of record - though prospective, multi-site, and peer-reviewed across 590,736 patients - was developer-led and observational, run by the party that built and commercialized the model, with no independent replication of the mortality effect. 'Peer-reviewed and prospective' is not 'independently validated'; the missing check is exactly the one developer-led evidence cannot supply about itself.

  • assumed

    No patient or clinical outcome is modeled here. This Lab reads institutional propagation only, and the patients being scored are not in the dynamics. The mortality benefit, its conditionality on confirmation, and the developer-led/observational caveats live in the case file, and are never computed from anything in this diagram.

What this example does not show

  • No patient or clinical outcome is modeled. The Lab reads institutional propagation only; the patients being scored are not in the dynamics, and the sepsis-mortality benefit, its conditionality on confirmation, and the developer-led/observational caveats live in the case file, never computed on this diagram.
  • The measured benefit is an observational association between provider confirmation and outcome, not a randomized effect of the algorithm; providers who confirmed alerts may differ systematically from those who did not, so confirmation may partly mark patients already likely to do better.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospectively across five hospitals of an academic health system covering 590,736 monitored patients — the largest prospective study of an ML sepsis system on record. Its central finding was conditional on the human loop: sepsis patients whose alert was evaluated and confirmed by a provider within three hours had a 3.3 percentage-point absolute and 18.7 percent relative adjusted reduction in in-hospital mortality, with less organ failure and shorter stays, while the alert on its own did not; a companion study found provider uptake varied with experience, unit culture, and alert context.

    empirical
    • Peer-reviewed Adams, R., Henry, K.E., et al. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28(7), 1455-1460. https://doi.org/10.1038/s41591-022-01894-0 https://www.nature.com/articles/s41591-022-01894-0
    • Peer-reviewed Henry, K.E., et al. (2022). Factors driving provider adoption of the TREWS machine learning-based early warning system and its effects on sepsis treatment timing. Nature Medicine, 28. https://doi.org/10.1038/s41591-022-01895-z https://www.nature.com/articles/s41591-022-01895-z
  • The TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: it was built at the deploying institution and commercialized through a company founded by its principal investigator, and confirmation-associated benefit is an observational association rather than a randomized effect of the algorithm — providers who engaged with alerts may differ from those who did not in ways the adjustment does not capture. The strongest numbers in the record therefore come from the party with the strongest interest in them, and no independent replication of the mortality effect had been published.

    empirical
    • Peer-reviewed Adams, R., Henry, K.E., et al. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28(7), 1455-1460. https://doi.org/10.1038/s41591-022-01894-0 https://www.nature.com/articles/s41591-022-01894-0

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical decision support & deterioration alerting domain page.

Levers available here and the patterns behind them

Documented case histories