Domain Atlas / Clinical decision support & deterioration alerting
IBM Watson for Oncology
Explore this deployment in the PAN Lab ↗
In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.
The least-established input behind this case's model organization's readings is an assumption, not a measurement. Evidence base: 3 assumed · 7 published baseline.
IBM's Watson for Oncology, a treatment-recommendation clinical decision support product trained by Memorial Sloan Kettering Cancer Center physicians, was in use at roughly 50 hospitals across five continents by September 2017, with only two named US adopters and primary markets in India, South Korea, China, Thailand and Mongolia; more than 70 Chinese institutions adopted it amid published reliability concerns, and some adopting hospitals marketed it to patients. It was sold as machine reading of the medical literature. It read no local health record: at each site a person abstracted the case by hand into a structured form of 13 to 17 attributes — 17 at a Danish pilot, 13 in Korean studies — and at Jupiter Medical Center in Florida a nurse spent roughly 90 minutes a week doing so, with the treating oncologist finding the output largely redundant with the plan already in hand. Its recommendation logic was undisclosed to buyers and IBM shipped multiple software versions without transparent change documentation, so evaluations run at different times were measuring different software.[4]
What happened
Two products carry the same technology name in this record, and conflating them is the single most common error in the secondary literature about it. Watson for Oncology (WFO) is the IBM product trained by Memorial Sloan Kettering Cancer Center physicians and sold to hospitals worldwide; it is the system the network in this case models. The MD Anderson Oncology Expert Advisor (OEA) is a different product — an MD Anderson-owned build on the same vendor technology — which was benched in September 2016 without ever reaching clinical use, was never piloted outside that institution, and is the subject of the University of Texas System procurement audit. The audit says nothing about Watson for Oncology. The internal slide decks that reported unsafe and incorrect recommendations say nothing about the Oncology Expert Advisor. This file keeps them apart in every paragraph.
Watson for Oncology ranked treatment options for named cancer types as recommended, for consideration, or not recommended. It was marketed as machine reading of the medical literature. What it actually computed over was a curated knowledge base built from a small number of SYNTHETIC cases — hypothetical patients, written by hand — authored by Memorial Sloan Kettering physicians working with IBM engineers, together with literature and guidelines those same physicians selected. STAT's July 2018 investigation, working from internal IBM documents, reported that the recommendations rested on the expertise of a few specialists for each cancer type rather than on guidelines or evidence; IEEE Spectrum's 2019 post-mortem records that synthetic cases were adopted as a workaround after the system could not learn from the literature. The buyers were not told this. The system also did not read local health records: at each site a person abstracted the case by hand into a structured form of 13 to 17 attributes — 17 at the Danish pilot, 13 in the Korean studies — and at Jupiter Medical Center in Florida a nurse spent roughly 90 minutes a week doing it, with the treating oncologist finding the output largely redundant with the plan already in hand. The recommendation logic was undisclosed, and IBM shipped multiple software versions without transparent change documentation, so evaluations run months apart were measuring different software with no way to say how it differed.
By September 2017 roughly 50 hospitals on five continents were using it, with only two named US adopters; the primary markets were India, South Korea, China, Thailand and Mongolia, and Tupasela and Di Nucci record more than 70 Chinese institutions adopting amid published reliability concerns. Some adopting hospitals marketed the system to patients. The evidence that accompanied all of this was concordance — how often clinicians chose what the system chose — and it is tiered, not uniform. The flagship study, 638 breast-cancer cases at Manipal in India published in Annals of Oncology in 2018, reported 93 percent agreement; that figure follows a blinded tumour-board re-review of the non-concordant cases which lifted agreement from 73 percent, it pools the recommended tier with the for-consideration tier, and its author list carries four IBM Watson Health-affiliated co-authors. The strongest independent evaluation, at Gachon Gil Medical Center in Korea with no IBM involvement and declared conflicts of none, measured agreement at the strict recommended level in 41.5 percent of 65 advanced gastric cancer patients (87.7 percent once the same pooling is applied) and named why the rest diverged: unaccounted patient history, US-centric regimens outdated or non-standard locally, national insurance not covering recommended agents, and the S-1 regimen — standard practice in Korea — simply absent from the MSK-derived knowledge base. A Danish pilot found roughly one-third agreement, its physician citing overweighting of American studies, and the hospital declined to adopt. No study anywhere measured a patient outcome. A peer-reviewed methodological critique in the journal AI and Society makes the structural objection: agreement is compatible with both parties being right and with both parties being wrong, disagreement cannot distinguish system error from clinician error, and the MSK-trained system was functioning as a de facto ground truth against which other countries' practice was being scored.
In June and July 2017 IBM Watson Health's deputy chief health officer presented internal slide decks to management which, in the documents STAT obtained and published a year later, recorded multiple examples of unsafe and incorrect treatment recommendations and stated that these raised serious questions about the process for building content and the underlying technology. That phrasing is IBM's own internal-document language as reported by STAT; IBM publicly contested STAT's characterisation, the examples came largely from testing and training exercises, and no patient harm from a recommendation appears anywhere in this record. What the record does show is what followed: global sales continued for years, the flagship Indian partner is reported by IEEE Spectrum to have discontinued its contract in December 2018, IBM halted Watson for Drug Discovery sales in 2019, and on 21 January 2022 IBM announced the sale of Watson Health's data and analytics assets to Francisco Partners. The deal closed on 30 June 2022, launching Merative around six named product families — Health Insights, MarketScan, Clinical Development, Social Program Management, Micromedex and Merge imaging — with the oncology treatment-recommendation products not among them. No formal discontinuation of Watson for Oncology was ever announced. No medical-device regulator reviewed its recommendations before it was marketed globally, and no regulator enforcement action, product withdrawal, or product-liability litigation over the recommendations was located as of August 2026.
The one formal accountability instrument in the whole arc that ever fired, fired on the twin. The University of Texas System Administration Audit Office's special review of the MD Anderson Oncology Expert Advisor, reported publicly in February 2017, found more than 62 million dollars paid to IBM and PricewaterhouseCoopers as of 31 August 2016 — approximately 39 to 40 million to IBM and 21 to 23 million to PwC across the contemporaneous accounts — an 11.59 million dollar deficit spend against donations not yet received, fees consistently set just below the amount that would have required Board of Regents approval, invoices paid in full regardless of whether contracted services were delivered, and standard IT-governance and competitive-procurement processes bypassed, with only one of seven reviewed service agreements competitively procured. IBM's support for the project ended 1 September 2016; IBM and the university agreed the system was not ready for human investigational or clinical use and that its use in the treatment of patients was prohibited, and it was not in clinical use and had not been piloted outside MD Anderson. Its proximate technical failure was infrastructure rather than oncology: built against MD Anderson's legacy ClinicStation record system, it was never updated to integrate with the Epic system the institution adopted in 2016, and its drug protocols and clinical-trial data had gone stale. The original June 2012 agreement had been extended twelve times. And the review, which is the source of every figure in this paragraph, expressly disclaimed any opinion on the scientific basis or functional capabilities of the system. The 62 million dollar figure is spend to date, not a fine and not an audit-assessed loss. The audit examined the money. Nothing anywhere examined the medicine.
The sociotechnical reading
The failure here is not a miscalibrated classifier — a model that sorts each case into ranked options — and not a compressed reviewer. It is the provenance of the knowledge store, and the network draws that as topology rather than as caption. The store is the model: one authored corpus, written by a few specialists at one institution, is what every ranked recommendation at every hospital on five continents was computed from. So the corpus write from those authors is the widest write on the diagram and it has no counterpart — the record contains no channel by which a hospital in Seoul, Copenhagen or Ulaanbaatar could add its own standard of care to the store its patients were being ranked against. The reconciliation edge that would carry deployed cases and their results back to the corpus is drawn at zero, because no such channel is documented and because no published study measured a patient outcome at all. What occupies that position instead is the opposite cargo: published agreement figures returning to the corpus as commercial validation of it. The Gachon investigators measured the consequence exactly — a chemotherapy regimen used routinely in their country, absent from the knowledge base, with no documented route by which the base could come to contain it.
Two further couplings follow from the first. The marketed capability is drawn as the pathway it would have been: the published literature is an input source whose direct read into the engine sits at zero, while the operator-mediated route — the authoring specialists selecting from that literature into the store, one judgment at a time — runs live. The distance between those two numbers is what the buyers were never told. And the deference gradient is drawn as two operator classes rather than as a sentence, because the record documents two groups meeting the same recommendation stream with opposite results. Subspecialist tumour boards had both the formal authority to overrule and the expertise that gives it force, and the record shows them using it: 41.5 percent agreement at the strict level in Korea, roughly one-third in Denmark followed by a decision not to buy. General-hospital clinicians had identical formal authority and no such expertise, and at one Mongolian hospital with no oncology specialists the recommendations were followed approximately 100 percent of the time. The advisory system was the decision-maker exactly where the capacity to catch its errors was lowest — and that is where the sales effort was growing, as US buyers held back. One asymmetry in the network is derived rather than assumed and is the sharpest thing on the diagram: every independent evaluation anywhere in this record was run at an institution that already had oncology subspecialty in the building. The deployments where the ranked list was followed most closely produced no evaluation, no published series and no measured disagreement of their own, so the pre-purchase check to that class is drawn at zero.
The governance reading is about the difference between measuring and controlling. The vendor's internal review is the clearest case in this atlas of a check that worked as measurement and did nothing as control: it did not merely list bad recommendations, it traced them to the process for building content — the right diagnosis, from a channel that was genuinely looking, reaching management in mid-2017 — after which the record shows years of continued selling, versions shipping without change notes, and an ending that arrived through a partner's contract exit and a business sale. Meanwhile the evaluative layer was captured by construction: with no outcomes channel anywhere, agreement with the tool became the currency, most of the favourable literature was produced by customers or with the vendor's own co-authors, the flagship design sent disagreeing human decisions back for re-review until agreement rose from 73 to 93 percent, and headline figures pooled two output tiers into one. The solver reads what all this means for governance and reports something worth stating plainly: the pathways this network cannot close at any budget are the writing of the corpus, the hand-keying of each case, the three writes that constitute every published study of the system, and the single-corpus coupling itself. Closing the first two would mean switching the deployment off; closing the third would mean publishing nothing about it; and there is no second corpus anywhere in the record to break the fourth against. The boundary holds as always: no patient outcome, treatment choice or injury is computed anywhere on this diagram, and none is documented anywhere in the record either — the absence of any outcome study is one of the case's findings rather than a gap in the model.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.