PAN Lab example
IBM Watson for Oncology
The authored corpus: an advisor sold as machine-read literature
A hospital buys a cancer treatment advisor sold as machine reading of the medical literature. A nurse keys thirteen to seventeen attributes from the chart into its form — the system reads no local record — and it returns treatment options ranked recommended, for consideration, or not recommended. Modelled on the oncology recommendation product IBM built with physicians at Memorial Sloan Kettering Cancer Center and sold to roughly fifty hospitals across five continents. What the ranking actually rests on is a knowledge base of synthetic cases: hypothetical patients written by a few specialists per cancer type at that one New York hospital, with the recommendation logic undisclosed and the provenance unlabelled to the buyer. So the store is the model, and one institution's practice is exported worldwide as ground truth. Two couplings follow. There is no return leg — no channel carries deployed cases, local standards or patient results back to the corpus, and the strongest independent study found a regimen used routinely in that country simply absent from it. And the same advice is safest where it is least needed: subspecialist tumour boards overruled it routinely, a Danish pilot measured roughly one third agreement and declined to buy, while a Mongolian hospital with no oncology specialists followed it approximately one hundred percent of the time. Agreement with the tool became the only evidence anyone published; no study measured a patient outcome. The vendor's own reviewers wrote the problem down in mid-2017 and the selling carried on. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets, and money is not what stops it. Inside the budget the failure regime can be brought to calm with the benefit reading clear of both margins. The gate that fails is the pathway gate, and it fails at any price — two levers costing six of your nine leave six pathways open, and the whole list at once at full strength, at more than four times your budget, leaves the same six. They are the writing of the corpus, the keying of each case into the form, the three writes that make up every published study of this system, and the single-corpus coupling itself. Closing the first two would mean switching the deployment off. Closing the third would mean publishing nothing about it. There is no second corpus to break the fourth against. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Watson-for-Oncology-class knowledge-provenance recommendation network: 11 components and 25 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 7 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
D48-derived new org (Phase 6, clinical-decision-support). TOPOLOGY. Eleven nodes and twenty-five pathways, all documented, none decorative: one model (a treatment-recommendation engine ranking options over a curated store and a keyed case form), three recordStores (the synthetic-case knowledge base, the structured case-attribute record, the concordance evaluation record — the three stores the PAN org draws), one inputSource (the published oncology literature and guidelines the product was marketed as reading), four operatorClasses and two reviewers, each reviewer with a documented inbound read. THE DRAWING IS AT THE COARSEST GRANULARITY THAT STILL HOLDS EVERY DOCUMENTED MECHANISM APART. Where the record documents one pathway that both treating classes walk — the read of the keyed abstraction, and the read of the published agreement literature — it is drawn once and narrated for both, because what the record measures differently about those two classes is the stream that reaches them and what each could do about it, not how they read a record they share. The ordinary workflow hand-off from the entry desk to the deciding room, the return of the ranked output through that same desk, and the authoring institution's reputation travelling with the product are all documented and all carried in the descriptions of the people they belong to; none is drawn as its own pathway, because the record shows none of them adopting or altering what passes through. FOUR ABSENCES ARE EQUALLY DERIVED. No enforcement node: the system was advisory end to end and no downstream action system is documented. No externalBoundary: no unsanctioned tool, ungoverned host or undocumented replication out of the governed system appears in this record, and the deployment's data-protection posture is undocumented rather than documented-bad. No worklist and no retriever: the documented burden is data-entry labour, roughly 90 nurse minutes per week at one named site, not a queue, and nothing in the build retrieves. No guardrail: the three-way output ranking is the shape of the output rather than a screen over it, and its documented consequence — published agreement figures pooling the recommended and for-consideration tiers — is carried on the study-record edges. THE MD ANDERSON ONCOLOGY EXPERT ADVISOR IS DELIBERATELY NOT DRAWN: it is a different product, it never reached clinical use, it was never piloted outside its own institution, and giving it nodes would assert a deployment the record refutes.
- baseline
THE DEFINING EDGE SET IS A WRITE MONOPOLY WITH NO RETURN LEG. The corpus write from the authoring physicians is drawn at 3 and has no counterpart from any deployed site, because none is documented: the record contains no channel by which a hospital in Seoul, Copenhagen or Ulaanbaatar could add its own standard of care to the store its patients were being ranked against. The reconciliation edge that would carry deployed cases and their outcomes back to the corpus is drawn AT 0 on the PAN org's own finding for this deployment — no channel returned real-world treatment results or local-standards corrections to the corpus — corroborated by the independent Korean evaluation, which measured a regimen standard in that country, absent from the corpus, and documented no route by which the corpus could come to contain it, and by the fact that no published study of this system measured patient outcomes anywhere. What the record does contain in that position is the opposite cargo: published agreement figures returning to the store as validation, drawn at 1 on the methodological critique's reading of the circularity.
- baseline
THE MARKETED CAPABILITY IS DRAWN AS A LATENT PATHWAY, NOT AS COPY. The product was sold as machine reading of the medical literature; the engineering post-mortem records that the system could not learn from the literature and that authoring synthetic cases was the workaround the programme adopted. So the literature is an inputSource whose direct read into the engine is drawn at 0, while the operator-mediated route — the authoring specialists selecting literature and guidelines into the store, one judgment at a time — is drawn at 2. The distance between those two numbers is what the buyers were not told: the investigation's finding that the recommendations rested on the expertise of a few specialists per cancer type rather than on guidelines or evidence is a statement about the ratio between the curated pathway and the authors' own synthetic vignettes.
- baseline
THE DEFERENCE GRADIENT IS DRAWN AS TWO OPERATOR CLASSES AND ONE CHECK AT 0. §5.1 licenses two classes where sources document distinct groups with different outcomes, and this record documents exactly that on the same recommendation stream. Subspecialist tumour boards: 41.5 percent agreement at the strict recommended level over 65 independent Korean cases, roughly one-third agreement at a Danish pilot that then declined to buy, a standing multidisciplinary board holding the decision at the Indian partner — inbound at 2. General-hospital treating clinicians: a Mongolian hospital with no oncology specialists following the recommendations approximately 100 percent of the time, more than 70 Chinese institutions adopting amid published reliability concerns, and hospitals in the international markets marketing the system to patients — inbound at 3. The second class writes nothing back anywhere in this network, and that is measured rather than assumed: every independent evaluation located in this record was run at a subspecialty-dense institution, and the deployments where the ranked list was followed most closely produced no evaluation, no published series and no measured disagreement of their own. Hence the pre-purchase check to that class is drawn AT 0. The two classes read the same keyed form and the same published literature, and those reads are drawn once each rather than twice, because no source measures either class reading either record differently; what the sources measure differently is the divergence from the ranked stream, which is where the two numbers sit. NOTE ON THE PAN SHAPE, IN FULL. The Lab network draws three things the PAN entry for this deployment folds rather than omits, and every one of them is the same evidence in a different vocabulary. (a) The PAN entry carries the treating side as ONE user class with the gradient held in that class's attributes and a population-weighted middle value; the Lab draws it as two classes, because the Lab's vocabulary can hold a gradient structurally and the PAN class attributes already name both ends of it. (b) The PAN entry carries independent site evaluation inside its purchasing-hospital governance actor rather than as a user; the Lab draws it as a reviewer, on the Danish pilot that evaluated and declined, the Korean study with no vendor involvement, and the methodological critique, all three of which the PAN entry cites. (c) The PAN entry has no separate literature entity, since the corpus store already carries curated literature among its contents; the Lab draws the published literature as an inputSource so that the marketed direct read and the actual operator-mediated route can be told apart as two pathways. Nothing is asserted anywhere here that the PAN file does not already record.
- baseline
DEMAND 2 / CAPACITY 1. Demand 2: this is a per-treatment-decision consult stream and no source documents an alert flood, a backlog or a waiting line anywhere in the deployment; the documented load is manual data-entry labour the system ADDS rather than absorbs — roughly 90 nurse minutes per week at one named US site, with the treating oncologist finding the output largely redundant with the plan already in hand — across roughly 50 hospitals on five continents with only two named US adopters. Capacity 1: the capacity that matters is the ability to catch a bad recommendation, and the record measures it as a function of local oncology subspecialty, which the sales footprint anticorrelated with. Where subspecialty was dense the check is documented working (a pilot that declined to adopt; an independent study measuring routine specialist divergence); where it was scarce the check inverted (approximately 100 percent adherence at a hospital without oncology specialists). A review capacity strongest exactly where it is least needed, on a footprint growing toward where it is weakest, reads 1.
- baseline
BASELINES. The corpus read into the engine is 3 because every ranked recommendation the product produced anywhere rested on that one store. The keyed case read is 2 — present on every case, bounded by an aperture of 13 to 17 fields that varied by site and by software version, with the independent evaluation tracing disagreement to what that aperture leaves out. The model self-loop is 3, the strongest documented single-corpus footprint in this catalogue: one authored knowledge base ranked into roughly 50 hospitals on five continents and more than 70 institutions in one country alone. The reviewer inbounds are 1 each: the vendor's review saw recommendation examples drawn largely from testing and training exercises, and independent site evaluation ran as case series at a handful of the adopting hospitals. Both reviewers' reads of the evaluation literature are 2, because reading that literature is what each of them demonstrably did well. What a buying hospital could look up in that same literature is narrated on the evaluation record itself rather than drawn as a pathway per treating class: the record documents who wrote the literature and what agreement can and cannot show, and it measures no reading of it by either class. The vendor's write into that literature is 1 on four named health-division co-authors of the flagship study and the company's own listing of it. The internal content-process review is 1 — the diagnosis is real and reached management, and the record after it is continued sales with no documented corpus correction.
- baseline
THE ATTRIBUTED FINDING IS ATTRIBUTED EVERYWHERE AND ASSERTED NOWHERE. The phrase recording multiple examples of unsafe and incorrect treatment recommendations is the vendor's own internal-document language of June and July 2017 as obtained and reported by a news organisation; the vendor publicly contested that characterization, the examples came largely from testing and training exercises, and no patient injury from a recommendation appears anywhere in this record. Every node and edge that touches the finding carries the attribution and the dispute, no parameter on this diagram is scaled by it, and this network never asserts harm to any patient. The portable-box framing belongs to the investigation and its named sources and is never presented as an admission by the authoring institution.
- baseline
THE PROCUREMENT AUDIT IS CARRIED IN COPY AND NOWHERE IN THE DIAGRAM, AND IT IS AN AUDIT OF THE MONEY. The University of Texas System special review, reported publicly in February 2017, examined the MD Anderson Oncology Expert Advisor — a different product, benched in September 2016 without clinical use and never piloted outside that institution — and found more than 62 million dollars paid to the vendor and a consulting firm as of 31 August 2016 (spend to date, not a loss finding), an 11.59 million dollar deficit spend against donations not yet received, fees consistently set just below the amount that would have required Board approval, invoices paid in full regardless of delivery, and procurement and IT governance bypassed, while expressly disclaiming any opinion on the scientific basis or functional capabilities of the system. It is the only formal accountability instrument anywhere in this arc that ever fired, it says nothing about the system this diagram models, and it is never cited here as a scientific verdict. No node, edge or lever represents it.
- assumed
WHERE THE RECORD IS SILENT, THE CONSERVATIVE VALUE. No source publishes a per-recommendation defect rate, an override rate for any class, a correction rate for the corpus, or record-hygiene measures for any of the three stores; the agreement figures are tiered evaluation results and are never treated as accuracy on this diagram. Those magnitudes are drawn within the qualitative rungs the documented figures support, matching the PAN org's own estimated-on-every-edge discipline — every edge in the PAN entry for this deployment is marked estimated, so drawn strengths here are modelling choices even where the pathway's presence and direction are evidence-derived.
- assumed
PATIENTS ARE NOT IN THE DYNAMICS. Cancer patients whose treatment options were ranked by this system, patients at hospitals that marketed it to them, and patients in specialist-scarce settings where the recommendations were followed nearly always are a boundary population recorded in the case file. No node, edge, baseline or lever here computes a treatment choice, an outcome or a harm for any patient, and none is documented anywhere in the record either: no published study of this system measured patient outcomes at all, which is one of the case's findings rather than a gap in this diagram. The PAN org's equity observations are deliberately empty for the same reason and this org adds none — the site-level differences in the record are properties of operator populations and deployment settings, drawn in the two treating classes where they belong.
What this example does not show
- Patients are not modelled. Nothing on this diagram computes a treatment choice, an outcome or a harm for any person, and nothing in the record does either: no published study of this system measured patient outcomes at all. No injury from any recommendation appears anywhere in this record, and this Lab asserts none.
- The concordance figures are tiered evaluation results, never accuracy. The 93 percent breast-cancer figure is from a study with four of the vendor's own health-division staff among its authors, pools the recommended and for-consideration output tiers, and follows a blinded re-review of the disagreeing cases that lifted agreement from 73 percent. The independent Korean figure at the strict recommended level is 41.5 percent over 65 patients, 87.7 percent once the same pooling is applied. A separate figure of roughly 49 percent for a Korean colon-cancer evaluation circulates in secondary accounts; it comes from a different study than the one parameterised here and belongs to the engineering post-mortem that reported it.
- This network models the globally sold oncology recommendation product only. The MD Anderson Oncology Expert Advisor was a different product, owned by that institution and built on the same vendor technology, which was benched in September 2016 without ever reaching clinical use and was never piloted outside its own hospital. Its University of Texas System special review — more than 62 million dollars paid to the vendor and a consulting firm as of 31 August 2016, an 11.59 million dollar deficit spend against donations not yet received, fees set just below the amount that would have required Board approval, procurement and IT governance bypassed — was a PROCUREMENT audit that expressly disclaimed any opinion on the system's scientific basis or functional capabilities. It says nothing about the system drawn here and is never cited as a scientific verdict. The 62 million dollar figure is spend to date, not a fine or a loss finding.
- The phrase recording multiple examples of unsafe and incorrect treatment recommendations is the vendor's own internal-document language of June and July 2017 as obtained and reported by a news organisation. The vendor publicly contested that characterisation, the examples came largely from testing and training exercises, and no patient harm is documented. The portable-box framing of the training institution's influence belongs to that investigation and its named sources, not to any admission by the institution.
- No litigation, regulator enforcement action or medical-device action over this system's recommendations was located as of August 2026. No medical-device regulator reviewed the recommendations before the product was marketed globally. The formal accountability record consists of the procurement review of the separate MD Anderson project plus market and corporate consequences — partner attrition, and the divestiture of the vendor's health business, announced 21 January 2022 and closed 30 June 2022, after which the successor company's named product families did not include the oncology recommendation products. No formal discontinuation of the product was ever announced, and this scenario asserts the end state only as those documented absences.
- The deployment's data-protection posture is undocumented rather than documented-bad. Identifiable clinical abstractions crossed from treating hospitals into a commercial system across many jurisdictions, but no breach, complaint, regulator finding or litigation over that handling appears anywhere in the record, and the privacy reading here is derived from the sensitivity of what flows rather than from any documented failure.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
IBM's Watson for Oncology, a treatment-recommendation clinical decision support product trained by Memorial Sloan Kettering Cancer Center physicians, was in use at roughly 50 hospitals across five continents by September 2017, with only two named US adopters and primary markets in India, South Korea, China, Thailand, and Mongolia; more than 70 Chinese institutions adopted it amid published reliability concerns, and some adopting hospitals marketed it to patients. It was sold as machine reading of the medical literature. It read no local health record: at each site a person abstracted the case by hand into a structured form of 13 to 17 attributes — 17 at a Danish pilot, 13 in Korean studies — and at Jupiter Medical Center in Florida a nurse spent roughly 90 minutes a week doing so, with the treating oncologist finding the output largely redundant with the plan already in hand. Its recommendation logic was undisclosed to buyers and IBM shipped multiple software versions without transparent change documentation, so evaluations run at different times were measuring different software.
empirical- Investigative Ross, C., & Swetlitz, I. (2017, September 5). IBM pitched its Watson supercomputer as a revolution in cancer care. It's nowhere close. STAT News https://www.statnews.com/2017/09/05/watson-ibm-cancer/
- Academic Tupasela, A., & Di Nucci, E. (2020). Concordance as evidence in the Watson for Oncology decision-support system. AI & SOCIETY https://link.springer.com/article/10.1007/s00146-020-00945-9
- Trade press Strickland, E. (2019, April). How IBM Watson Overpromised and Underdelivered on AI Health Care. IEEE Spectrum https://spectrum.ieee.org/how-ibm-watson-overpromised-and-underdelivered-on-ai-health-care
- Academic JNCI: Journal of the National Cancer Institute (2017). M.D. Anderson Breaks With IBM Watson, Raising Questions About Artificial Intelligence in Oncology, 109(5) https://academic.oup.com/jnci/article/109/5/djx113/3847623
The system's recommendations were computed from a curated knowledge base built on a small number of SYNTHETIC — hypothetical, hand-written — cancer cases authored by Memorial Sloan Kettering physicians working with IBM engineers, together with literature and guidelines those physicians selected, resting on the expertise of a few specialists for each cancer type rather than on guidelines or evidence; the engineering post-mortem records that synthetic cases were adopted as a workaround after the system could not learn from the medical literature. Internal IBM Watson Health slide decks of June and July 2017, presented by the division's deputy chief health officer to management and later obtained by STAT, recorded 'multiple examples of unsafe and incorrect treatment recommendations' and stated that these raised 'serious questions about the process for building content and the underlying technology.' That phrasing is IBM's own internal-document language as reported by STAT; IBM publicly contested STAT's characterisation, the examples came largely from testing and training exercises, and no patient harm from a recommendation is documented anywhere in this record. Global sales continued for years afterwards, and no correction of the knowledge base following the finding has been located.
empirical- Investigative Ross, C., & Swetlitz, I. (2018, July 25). IBM's Watson supercomputer recommended 'unsafe and incorrect' cancer treatments, internal documents show. STAT News https://www.statnews.com/2018/07/25/ibm-watson-recommended-unsafe-incorrect-treatments/
- Trade press Strickland, E. (2019, April). How IBM Watson Overpromised and Underdelivered on AI Health Care. IEEE Spectrum https://spectrum.ieee.org/how-ibm-watson-overpromised-and-underdelivered-on-ai-health-care
- Investigative Ross, C., & Swetlitz, I. (2017, September 5). IBM pitched its Watson supercomputer as a revolution in cancer care. It's nowhere close. STAT News https://www.statnews.com/2017/09/05/watson-ibm-cancer/
No published study measured whether Watson for Oncology improved patient outcomes; the published evidence base is concordance — how often clinicians chose what the system chose — and it is tiered rather than uniform. The flagship study of 638 breast-cancer cases at Manipal, published in Annals of Oncology in 2018, reported 93 percent agreement, a figure reached after a blinded tumour-board re-review of the non-concordant cases lifted agreement from 73 percent, pooling the 'recommended' and 'for consideration' output tiers, on an author list carrying four IBM Watson Health-affiliated co-authors. The strongest independent evaluation, at Gachon Gil Medical Center in Korea with no IBM involvement and declared conflicts of none, measured agreement at the strict 'recommended' level in 41.5 percent of 65 advanced gastric cancer patients (87.7 percent counting 'for consideration'), attributing the divergence to unaccounted patient history, US-centric regimens outdated or non-standard locally, national insurance not covering recommended agents, and the S-1 regimen — standard practice in Korea — being absent from the knowledge base. A Danish pilot found roughly one-third agreement, its physician citing overweighting of American studies, and the hospital declined to adopt; at UB Songdo Hospital in Mongolia, which lacked oncology specialists, clinicians followed the recommendations approximately 100 percent of the time. A peer-reviewed methodological critique concludes that concordance cannot establish safety or benefit — agreement is compatible with both parties being wrong, and disagreement cannot distinguish system error from clinician error — and documents the MSK-trained system functioning as a de facto ground truth against which other countries' practice was scored.
empirical- Academic Somashekhar, S. P., et al. (2018). Watson for Oncology and breast cancer treatment recommendations: agreement with an expert multidisciplinary tumor board. Annals of Oncology, 29(2), 418-423 (IBM Watson Health co-authors) https://academic.oup.com/annonc/article-abstract/29/2/418/4781689
- Academic Canadian Journal of Gastroenterology and Hepatology (2019). Concordance Rate between Clinicians and Watson for Oncology among Patients with Advanced Gastric Cancer: Early, Real-World Experience in Korea (Gachon Gil Medical Center) https://pmc.ncbi.nlm.nih.gov/articles/PMC6377977/
- Academic Tupasela, A., & Di Nucci, E. (2020). Concordance as evidence in the Watson for Oncology decision-support system. AI & SOCIETY https://link.springer.com/article/10.1007/s00146-020-00945-9
- Investigative Ross, C., & Swetlitz, I. (2017, September 5). IBM pitched its Watson supercomputer as a revolution in cancer care. It's nowhere close. STAT News https://www.statnews.com/2017/09/05/watson-ibm-cancer/
The MD Anderson Oncology Expert Advisor was a separate product — an MD Anderson-owned build on IBM Watson technology, distinct from Watson for Oncology — which never reached clinical use. IBM support ended 1 September 2016; IBM and the university agreed the system was 'not ready for human investigational or clinical use, and its use in the treatment of patients is prohibited,' and it was 'not in clinical use and has not been piloted outside of MD Anderson.' The University of Texas System Administration Audit Office's special review, reported publicly in February 2017, found more than 62 million dollars paid to IBM and PricewaterhouseCoopers as of 31 August 2016 (approximately 39 to 40 million to IBM and 21 to 23 million to PwC across contemporaneous accounts), an 11.59 million dollar deficit spend against donations not yet received, fees consistently set just below the amount that would have required Board approval, invoices paid regardless of delivery, and procurement and IT-governance processes bypassed — while expressly disclaiming any opinion on the scientific basis or functional capabilities of the system. The figure is spend to date, not a fine or an audit-assessed loss, and the review is a procurement audit, not a scientific verdict on either product. No regulator enforcement action, medical-device action, or product-liability litigation over either system's recommendations was located as of August 2026, and no medical-device regulator reviewed Watson for Oncology's recommendations before it was marketed globally. IBM announced the sale of Watson Health's data and analytics assets to Francisco Partners on 21 January 2022; the sale closed on 30 June 2022, launching Merative around six named product families with the oncology treatment-recommendation products not among them. No formal discontinuation of Watson for Oncology was announced.
empirical- Government University of Texas System Administration Audit Office (2016, November; public February 2017). Special Review of Procurement Procedures Related to the M.D. Anderson Cancer Center Oncology Expert Advisor Project https://www.utsystem.edu/sites/utsfiles/documents/system-audit/ut-system-administration-special-review-procurement-procedures-related-utmdacc-oncology-expert-advis/ut-system-administration-special-review-procurement-procedures-related-utmdacc-oncology-expert-advis.pdf
- Trade press The Register (2017, February 20). Watson can't cure cancer ... or all the stuff that breaks IT projects; and PCWorld / IDG News Service (2017). Texas hospital struggles to make IBM's Watson cure cancer https://www.theregister.com/2017/02/20/watson_cancerbusting_trial_on_hold_after_damning_audit_report/
- Academic JNCI: Journal of the National Cancer Institute (2017). M.D. Anderson Breaks With IBM Watson, Raising Questions About Artificial Intelligence in Oncology, 109(5) https://academic.oup.com/jnci/article/109/5/djx113/3847623
- Vendor IBM Newsroom (2022, January 21). Francisco Partners to Acquire IBM's Healthcare Data and Analytics Assets; and Francisco Partners via Business Wire (2022, June 30). Francisco Partners Completes Acquisition of IBM's Healthcare Data and Analytics Assets; Launches Healthcare Data Company Merative https://newsroom.ibm.com/2022-01-21-Francisco-Partners-to-Acquire-IBMs-Healthcare-Data-and-Analytics-Assets
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
All of them in context on the Clinical decision support & deterioration alerting domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Gate vendor updates — Vendor quality gate
- Mark AI-written records — Provenance labeling
- Check copied records — Reconcile copied records
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Assign a challenger — Structured dissent
- Pause AI on alarms — Deployment circuit-breaker
Documented case histories
- IBM Watson for Oncology
- TREWS sepsis early-warning system
- Advance Alert Monitor (AAM) deterioration model
- Sepsis Watch deep-learning detection system
- Proprietary EHR sepsis model (external validation)
- nH Predict Utilization Review
- Cost-Proxy Care Stratification
- CA-CDS Child Abuse Alerting
- IDx-DR Autonomous Screening
- Viz.ai LVO Stroke Triage
- OPTN eGFR Waiting-Time Correction
- Practice Fusion Pain CDS
- UBH Level of Care Guidelines (Wit v. UBH)
- EviCore by Evernorth: the review threshold
- Cigna PxDx