Skip to content

PAN Lab example

Woebot (a governed app wind-down)

The responsible wind-down: retiring a peer-reviewed CBT chatbot

A rule-based CBT chatbot ran for years as a self-help app, used by roughly 1.5 million people over its lifetime, and its maker chose to retire it — deliberately, on a published schedule, with a window to download your transcripts and a date by which all account data would be anonymized. Modeled on Woebot. The tool was non-generative (authored decision paths, not free-form text) and its foundational peer-reviewed study — authored by its own maker's team — was a small, early-stage efficacy signal, so the danger here is not a runaway model error. It is what a running service does with the sensitive record it accumulated when the business reason to keep going runs out, and whether an early-stage signal gets quietly read as validation on the way out. Watch the transcript store, and watch what the exit is honest about.

Stylized model of a documented deploymentBehavioral-health & crisis triage

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Woebot-class governed wind-down of a scripted CBT chatbot network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the governed-wind-down pattern documented in the Woebot case file — not a reconstruction of the actual app or its scripted content. The tool modeled here is the rule-based, non-generative direct-to-consumer and enterprise cognitive behavioural therapy (CBT) chatbot, not the separate investigational prescription variant (WB001), and the two are never conflated. Two modeling bounds are disclosed: the session-context read is an assumed pathway (the documented record covers storage, user retrieval, and scheduled anonymization only), and the documented enterprise and health-system distribution partners — institutional counterparties a governed wind-down must notify and disentangle — are not drawn as nodes; the map models the maker's own institutional loop.

  • assumed

    This is a positive-pole, safe-shape network, like the verify-before-use copilot: the model is non-generative (authored decision paths, not free-form generation), so its per-interaction error surface is bounded by design; the operator loop is institutional monitoring rather than a stretched operator adopting outputs; and the intake is voluntary self-report collected for this purpose. Contamination is low by design at baseline, and the question flips to what a responsible exit does with the record it accumulated.

  • baseline

    The shape's defining feature is the teardown of the record store. The per-user transcript store is a large corpus of the most sensitive possible content, and the governed exit is what keeps it from leaking or being silently abandoned: a transcript-request window, then scheduled anonymization removing personally identifying information. The record's pathway out of the org is drawn empty at baseline because that is the pathway a bad wind-down opens and a governed one closes. data-minimization (bounding what is kept and how identifiable it is) and connection-auth (structurally closing the pathway out) act on it; a content-blind record-purge does not close it and backfires — the responsible teardown minimizes and anonymizes by content rather than deleting blind.

  • baseline

    A second absence is on the evidence side, drawn as the empty independent model check: no independent, arm's-length evaluation is documented to have checked what the tool's evidence actually showed. The foundational peer-reviewed study was authored by the tool's own personnel, and it is a small (n=70), short (two weeks), unblinded, early-stage efficacy signal, not regulatory validation; the tool's broader published evidence base is not assembled here, and the absence claimed is of an independent evaluation, not of further studies. The consumer app itself was never a Food and Drug Administration (FDA)-authorized device, and the investigational prescription variant never received FDA marketing authorization. oversight-cadence, conformity-gate, and provenance-labels open this edge.

  • assumed

    The terminal event in the real case was exogenous — the founder attributed the shutdown to the cost of Food and Drug Administration (FDA) marketing-authorization and a regulatory-pathway gap, framed as economic and regulatory rather than a clinical failure (her own on-record account, not an independently audited finding). On the map, that means the danger is not a runaway model error; it is what happens to the accumulated record store when a running service is decommissioned, and whether the exit is executed on a published schedule or left to lapse under decommissioning pressure.

  • assumed

    The regulatory-and-clinical oversight node carries a real, lighter review pathway of its own — Food and Drug Administration (FDA) engagement for the investigational prescription variant plus internal clinical leadership above the consumer app — so it enters the dynamics rather than sitting on a pathway.

  • assumed

    Served people (the app's end users, roughly 1.5 million over the tool's lifetime, a cumulative figure) are not in the dynamics. No clinical, symptom, or crisis outcome is computed here; this Lab models institutional propagation only, and a conversation, a transcript, or a routing on this map is an institutional signal, never a person. The operator network modeled here is Woebot Health's own clinical and operations staff and, in the prescription pathway, supervising clinicians.

What this example does not show

  • This Lab models institutional propagation only. It never models symptoms, recovery, crisis, suicide, or any clinical outcome, and the people who used this app are not in the diagram — a conversation, a transcript, or a routing here is an institutional signal, never a person. The tool's clinical value and any effect on the people it served are documented in the case file and measured outside any diagram like this one.
  • The efficacy record is hedged as the sources hedge it: the foundational, vendor-authored 2017 randomized controlled trial (n=70, ages 18 to 28, two weeks, unblinded, information-only control) reported a moderate between-groups reduction in depression symptoms (about d equals 0.44). That is an early-stage efficacy signal, not regulatory validation or a generalizable effectiveness claim, and it is carried here as such throughout; the tool's broader published evidence base is not assembled here, and what the record lacks is an independent, arm's-length evaluation, not further vendor-side studies.
  • The two products are kept distinct. The tool modeled here is the rule-based, non-generative direct-to-consumer and enterprise app that was retired. The investigational, prescription-only variant (WB001) received a Food and Drug Administration (FDA) Breakthrough Device Designation — an expedited-review status, not marketing authorization — and entered a pivotal medical-device trial, but never received FDA marketing authorization; the two are never conflated.
  • The cause of the shutdown is presented as reported, not as an independently audited fact: the founder attributed it to the cost of Food and Drug Administration (FDA) marketing-authorization and a regulatory-pathway gap, framing the exit as economic and regulatory rather than a clinical failure. That is a self-reported account. The lifetime user figure (roughly 1.5 million) is a cumulative number reported in press coverage, not an audited point-in-time active-user count, and funding figures differ across sources (the primary press releases give a 90M Series B bringing total funding to 114M; a secondary summary labels it a Series C and totals about 107.5M).

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Woebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people over its lifetime, was deliberately retired by its maker on a pre-announced schedule: the app was taken down on June 30, 2025, with a transcript-request window (deadline July 15, 2025) and all account data anonymized as of July 31, 2025, removing personally identifying information rather than silently abandoning the service. The founder and chief executive attributed the shutdown to the cost of meeting FDA marketing-authorization requirements and to a regulatory-pathway gap, framing the exit as economic and regulatory rather than a clinical failure - a self-reported account, not an independently audited finding. The roughly 1.5 million figure is a cumulative lifetime number reported in press coverage, not an audited point-in-time active-user count.

    empirical
    • Investigative Aguilar, Why Woebot, a pioneering therapy chatbot, shut down (STAT News, 2025) https://www.statnews.com/2025/07/02/woebot-therapy-chatbot-shuts-down-founder-says-ai-moving-faster-than-regulators/
    • Vendor Woebot Health, FAQs (Woebot app retirement) (2025) https://woebothealth.com/faq/
    • Trade press HLTH, Woebot Health Is Shutting Down Its App (2025) https://hlth.com/insights/news/woebot-health-is-shutting-down-its-app-2025-04-28
  • The foundational study in Woebot's peer-reviewed efficacy record is an early-stage, vendor-authored trial: a 2017 randomized controlled trial in JMIR Mental Health (n=70, ages 18 to 28, two weeks, unblinded, information-only control) reported a moderate between-groups reduction in PHQ-9 depression symptoms (about d = 0.44). That is an efficacy signal, not regulatory validation; the study authors were affiliated with the tool's maker, and no independent, arm's-length evaluation of the consumer app is documented (the broader published evidence base is not assembled here). A separate, investigational, prescription-only variant (WB001) received an FDA Breakthrough Device Designation in May 2021 - an expedited-review status, not marketing authorization - and entered a pivotal Software as a Medical Device trial with the first patient enrolled in January 2023, but never received FDA marketing authorization; it must not be conflated with the consumer app.

    empirical
    • Academic Fitzpatrick, Darcy, Vierhile, Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial (JMIR Mental Health, 2017;4(2):e19) https://mental.jmir.org/2017/2/e19/
    • Vendor Woebot Health (Business Wire), Woebot Health Receives FDA Breakthrough Device Designation for Postpartum Depression Treatment (2021) https://www.businesswire.com/news/home/20210526005054/en/Woebot-Health-Receives-FDA-Breakthrough-Device-Designation-for-Postpartum-Depression-Treatment
    • Vendor Woebot Health (Business Wire), Woebot Health Enrolls First Patient in Pivotal Clinical Trial of WB001 for Postpartum Depression (2023) https://www.businesswire.com/news/home/20230123005211/en/Woebot-Health-Enrolls-First-Patient-in-Pivotal-Clinical-Trial-of-WB001-for-Postpartum-Depression

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Behavioral-health & crisis triage domain page.

Levers available here and the patterns behind them

Documented case histories