PAN Lab example
Tessa chatbot replacing the NEDA eating-disorder helpline
The warrant stayed put while the system moved
A national nonprofit replaced a twenty-year helpline, which had fielded nearly 70,000 contacts in a year, with a chatbot. Modeled on a deployment whose tool had two layers with different evidence behind them: a closed, pre-scripted prevention programme with a published randomised trial, and a question-and-answer feature the operating vendor added later. Answers the trial never covered went out under the name the trial warranted. The tool was pulled two days before it was to become the sole channel, and the helpline closed as scheduled anyway. So watch two seams: who is allowed to change what the system can say, and whether the evidence it is presented on still describes it. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. Cost is not what blocks it. Inside the budget the best legal settings clear the benefit margins comfortably. Lift the pathway requirement on its own and the board wins at both tiers, from a stack costing 4 of the 13 you have under Service and Safety Targets and 6 under All Governance Targets. The pathway gate is the only gate that fails. One pathway stays open at every affordable price: the vendor changing what the answering layer is allowed to say. That is the authority the record leaves unresolved, and nothing the client holds closes it. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Helpline-replacement-class where the evidence and the artifact came apart network: 10 components and 17 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 7 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
The two model nodes are the derivation's first and largest choice, and they come straight from the record: the deployment is documented as a two-layer artifact whose layers had different evidence status - a closed, pre-scripted prevention programme with a published randomised trial behind it, and a generative question-and-answer feature the operating vendor added afterwards. Drawing one model would have made the case unreadable, because the case is the coupling: answers the trial never covered went out through the channel the trial warranted. The strongest edge on the board is therefore the model-to-model link at full strength, and the strongest inbound is the vendor's capability change at full strength, which is the documented inversion - discretion over what the system could say sat with the vendor, and none sat at the point of service.
- assumed
This models the helpline-replacement pattern documented in the case file - not a reconstruction of the actual chatbot or its content. Recency stamp, stated up front per the case's binding framing: this is a closed 2022-2023 arc used as a documented archetype, not as evidence about present-day vendor practice. Its unusual value is that every phase is documented - trial baseline, deployment, capability change, harm discovery, suspension - which almost no open case offers.
- baseline
Three contested channels are kept contested, and none of them is resolved on this diagram. The vendor's claim that the added feature was covered by the contract and the client's denial are both claims: this is modelled as contested authority on the vendor-to-model edge, never as an established breach. The attribution of the May 2023 dieting advice to the generative layer is the reported and dominant explanation, and the vendor's own public statements were mixed, so the edge copy states it as attributed rather than as established mechanism. Anti-union motive is asserted by the workers and the labor press and denied by the organisation; the four-day gap between certification and the termination announcement is documented fact, the motive is not, the unfair-labor-practice outcome is not publicly documented, and no finding of retaliation is asserted anywhere here.
- baseline
The record node is the programme's published warrant, not a case file, and the two writes into it are the point. The closed programme's tested behaviour is what the published trial documents, drawn faint as a bounded write. What the deployment actually produced after the answering feature was added has no documented write into that record at all, so that pathway is drawn empty - the documented sentence is that the evidence record was pinned to the rule-based artifact and never updated for the generative variant. Meanwhile the two reads off the record run strong, because the service was presented publicly on the trial's name and evidence. A record that keeps asserting a warrant the shipped system no longer matches is this deployment's signature, and provenance labelling is the lever that acts on exactly those reads.
- baseline
The external clinical report is drawn faint, not empty, and that is the case's binding corrected reading. The earliest documented warning, from a peer organisation's executive director in October 2022, did produce a change: the specific language she flagged was removed soon after. What did not follow was any documented systemic review, output monitoring, or reassessment of the deployment, and system-level action - suspension - came about seven months later and was triggered by public screenshots rather than by the report. The provenance of the flagged content is itself contested: the vendor said it was pre-scripted, the research team denies writing it, so this warning is the earliest documented harmful-output warning and cannot be read as an early alarm about the generative layer.
- baseline
Three operator classes are drawn because the record documents three different groups with different roles and different outcomes: the helpline associates, supervisors and volunteers who ran the incumbent service; the vendor operations team that hosted the deployment and changed its capability; and the communications staff who were the only humans documented as acting on the system's output after the transition. The empty middle is the finding - the record states there was no human in the conversation loop and no real-time monitor - and drawing the three groups that did exist is what makes that emptiness visible without inventing a node to represent it.
- baseline
The queue is drawn as a tethered mediator carrying no flow, and its figures are labelled for what they are. The deploying organisation reported that 46% of initial contacts could not be answered immediately and that message replies lagged six to eleven days; those numbers originate with the organisation itself, were relayed by journalists, and formed part of its public case for the change, so they are carried as the justification narrative they are and are never treated as an independent measurement of the human service's capacity. That is also why the manual capacity is drawn high rather than marked down: a twenty-year service handling nearly 70,000 contacts a year with referral and de-escalation is a real counterfactual floor, and the engine must be able to read the replacement as net-negative against it.
- baseline
The egress to the vendor-hosted environment is drawn live rather than empty because the hosting is documented as actual, not hypothetical: the vendor operated and hosted the deployment, the added feature's training data is undisclosed in the public record, its safety claim was not independently verified, and the unilateral capability change is itself the proof that the client's control over that environment did not hold. It is marked privacy-sensitive because the payload is free-text help-seeking about eating disorders. No data-protection breach, sale, or secondary use is documented, and none is implied here.
- baseline
The published trial is a baseline for the closed programme only. It tested prevention in 700 screened women at high risk for an eating disorder against a waiting-list arm, reporting small effects on weight and shape concerns at three and six months, an overall psychopathology effect at three months that was not sustained at six, better odds of remaining non-clinical at both points, and no effect on depression or anxiety. It is not efficacy evidence for the generative variant and not efficacy evidence for replacing a helpline, and nothing on this diagram may be read as either.
- assumed
Served people are not in the dynamics and no clinical outcome is computed here. The people who typed into the widget or called the line are boundary-only; this Lab reads institutional propagation, and a contact, a reply, or a referral on this map is an institutional signal, never a person. The harmful-output content, the differential risk of dieting advice to a population with eating disorders, and the capacity gap left when both channels were offline live in the case file and are never derived from anything on this diagram. Prevalence of harmful replies in ordinary use is unknown: the documented instances come from adversarial testers and a peer-organisation director, not from routine-user logs.
What this example does not show
- Served people are not modeled. The Lab reads institutional propagation only, so the people who typed into the widget or called the line are boundary-only, and no clinical, eating-disorder, or crisis outcome is computed from anything on this diagram. The harmful-output content, the differential risk that ordinary dieting advice carries for this population, and the capacity gap left when both channels were offline live in the case file, never derived here.
- This is a closed 2022-2023 arc used as a documented archetype, not as evidence about present-day vendor practice. Its value is that every phase is documented - trial baseline, deployment, capability change, harm discovery, suspension - which almost no open case offers, and the recency exception is stated rather than assumed.
- Three channels stay contested. The vendor's claim that the added feature was covered by the contract and the client's denial are both claims, modeled as contested authority with no adjudicated breach. The attribution of the May 2023 dieting advice to the generative layer is the reported and dominant explanation, and the vendor's own public statements were mixed, so it is entered as attributed rather than established mechanism. Anti-union motive is asserted by the workers and the labor press and denied by the organisation: the four-day gap between certification and the termination announcement is documented fact, the motive is not, the unfair-labor-practice outcome is not publicly documented, and no finding of retaliation is asserted here.
- The earliest documented warning did produce a change. The specific language flagged in October 2022 was removed soon after, and what is absent from the record is a systemic review, output monitoring, or reassessment of the deployment. The provenance of that flagged content is itself contested - the vendor said it was pre-scripted, the research team denies writing it - so it is the earliest documented harmful-output warning and not an early alarm about the generative layer.
- The demand and queue figures are the deploying organisation's own. Nearly 70,000 contacts, growth of more than 100% over pre-pandemic levels, 46% of initial contacts not answered immediately, and a six-to-eleven-day message lag all originate with the organisation, were relayed by journalists, and formed part of its public case for the change. They are drawn as the justification narrative they are, and they are not treated as an independent measurement of the human service they were used to argue against.
- The published trial is a baseline for the closed programme only. It tested prevention in screened at-risk women against a waiting list and is not efficacy evidence for the deployed replacement or for the generative variant. Prevalence of harmful replies in ordinary use is unknown: every documented instance came from adversarial testers and a peer-organisation director, not from routine-user logs.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Four days after a labor board certified its paid helpline staff's union election, a national eating-disorder nonprofit told those staff they were terminated and that a chatbot would replace a helpline that had fielded nearly 70,000 contacts in 2022 with six paid staff, about two supervisors, and up to roughly 200 trained volunteers; the chatbot was suspended on May 30, 2023, two days before it was to become the sole channel, after testers published screenshots of it recommending a 500 to 1,000 calorie daily deficit, a 1 to 2 pound weekly loss, a 2,000 calorie cap, regular weigh-ins, and where to buy skinfold calipers, and the helpline closed on June 1 as scheduled while the chatbot was already offline.
empirical- Investigative Wells, Can a chatbot help people with eating disorders as well as another human? (NPR / Michigan Public, 2023) https://www.npr.org/2023/05/24/1177847298/can-a-chatbot-help-people-with-eating-disorders-as-well-as-another-human
- Investigative KFF Health News, What Does a Chatbot Know About Eating Disorders? Users of a Help Line Are About to Find Out (2023) https://kffhealthnews.org/news/article/what-does-a-chatbot-know-about-eating-disorders-users-of-a-help-line-are-about-to-find-out/
- Investigative NPR, An eating disorders chatbot offered dieting advice, raising fears about AI in health (2023) https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea
- Investigative Picchi, Eating disorder helpline shuts down AI chatbot that gave bad advice (CBS News, 2023) https://www.cbsnews.com/news/eating-disorder-helpline-chatbot-disabled/
- Investigative NPR, National Eating Disorders Association phases out human helpline, pivots to chatbot (2023) https://www.npr.org/sections/health-shots/2023/05/31/1179244569/national-eating-disorders-association-phases-out-human-helpline-pivots-to-chatbo
The deployed chatbot had two layers with different evidence status: a closed, pre-scripted program the vendor and the research team describe as unable to depart from its authored content, and a generative question-and-answer feature the operating vendor added in what its chief executive called a systems upgrade covered by the client's contract, a reading the client's chief executive denies by saying the organization was never advised of the changes and would not have approved them; the vendor's public account of the harmful outputs was mixed, saying it was still trying to determine how a closed system allowed such content, so the attribution of the May 2023 advice to the generative layer is the reported and attributed explanation rather than an established mechanism, and the contract scope stays contested with no adjudicated breach.
empirical- Investigative NPR, An eating disorders chatbot offered dieting advice, raising fears about AI in health (2023) https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea
- Investigative KFF Health News, What Does a Chatbot Know About Eating Disorders? Users of a Help Line Are About to Find Out (2023) https://kffhealthnews.org/news/article/what-does-a-chatbot-know-about-eating-disorders-users-of-a-help-line-are-about-to-find-out/
- Investigative TheWrap (republished on Yahoo), Eating Disorder Chatbot Taken Down After Giving Weight Loss Advice, Nonprofit Blames Bad Actors (2023) https://www.yahoo.com/entertainment/eating-disorder-chatbot-taken-down-164509547.html
The earliest documented external warning about harmful chatbot responses came in October 2022 from the executive director of a peer eating-disorder organization, and the specific language she flagged was quickly removed after she reported it, with no documented systemic review, output monitoring, or reassessment of the deployment following; the vendor's chief executive said the flagged language was part of the pre-scripted content rather than the generative layer, which the research team denies, so its provenance is contested, and system-level action arrived roughly seven months later when public screenshots circulated.
empirical- Investigative NPR, An eating disorders chatbot offered dieting advice, raising fears about AI in health (2023) https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Assign a challenger — Structured dissent
- Peer sharing rules — Peer-edge governance
- Escalate checks — State-feedback vigilance
- Check with a second model — Cross-model verification
- Pause AI on alarms — Deployment circuit-breaker
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- Tessa chatbot replacing the NEDA eating-disorder helpline
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Kaiser Permanente Suicide-Risk Model
- Crisis Text Line & Loris.ai
- LyssnCrisis counselor QA at ProtoCall Services (988)
- NarxCare
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- Woebot (a governed app wind-down)