ParamergeParamerge

Domain Atlas / Behavioral-health & crisis triage

Case fileUnited States (national nonprofit; helpline staff based in New York City; the only formal external proceedings were before the National Labor Relations Board, on the labor channel)medium deployment

Tessa chatbot replacing the NEDA eating-disorder helpline

Four days after a labor board certified its paid helpline staff's union election, a national eating-disorder nonprofit told those staff they were terminated and that a chatbot would replace a helpline that had fielded nearly 70,000 contacts in 2022 with six paid staff, about two supervisors, and up to roughly 200 trained volunteers; the chatbot was suspended on May 30, 2023, two days before it was to become the sole channel, after testers published screenshots of it recommending a 500 to 1,000 calorie daily deficit, a 1 to 2 pound weekly loss, a 2,000 calorie cap, regular weigh-ins, and where to buy skinfold calipers, and the helpline closed on June 1 as scheduled while the chatbot was already offline.[5]

What happened

Recency stamp first, because it governs how everything below should be read. This is a closed 2022-2023 arc, used here as a documented archetype and not as evidence about present-day vendor practice. It earns its place because almost no open case offers what this one does: a published trial baseline, a documented deployment, a documented capability change, a documented harm discovery, and a documented suspension, all in the public record.

Tessa was a chatbot built by researchers at Washington University School of Medicine with Stanford collaborators, funded by a 2018 grant from the National Eating Disorders Association. It delivered the "Body Positive" program, an adaptation of the StudentBodies cognitive-behavioral prevention curriculum, as a closed, pre-scripted system. Its creator, Ellen Fitzsimmons-Craft, described the design constraint plainly: "By design it couldn't go off the rails... AI isn't ready for this population." A randomized clinical trial published in the International Journal of Eating Disorders enrolled 700 women at high risk for an eating disorder and compared the chatbot against a waitlist. It reported small effects on weight and shape concerns, d = -0.20 at three months and d = -0.19 at six; an effect on overall eating-disorder psychopathology of d = -0.29 at three months that was not sustained at six; better odds of remaining non-clinical, OR 2.37 (95% CI 1.37-4.11) at three months and OR 2.13 (1.26-3.59) at six; and no effect on depression or anxiety. That is a real result for prevention in a screened at-risk college population, and it is not efficacy evidence for replacing a helpline or for anything the tool did after it changed.

Tessa went live on the organization's website in February 2022, alongside the human helpline. At some point in 2022 the operating vendor, Cass (formerly X2AI), added a generative "enhanced question-and-answer feature." Its chief executive, Michiel Rauws, described this as a "systems upgrade" that was "part of NEDA's contract." The organization's chief executive, Liz Thompson, said "NEDA was never advised of these changes and did not and would not have approved them," and the research team said the generative capability was outside their design. That contract dispute is unresolved in the public record, and nothing here treats either account as established.

The labor channel ran on its own clock. The paid helpline staff, who had asked for adequate staffing and training rather than raises, won a National Labor Relations Board election on March 17, 2023 as Helpline Associates United with Communications Workers of America Local 1101. Certification came on March 27. On March 31, four days later, board chair Geoff Craddock told all paid helpline staff on a call NPR later obtained audio of that they were terminated and that Tessa would replace the helpline effective June 1. The organization's vice president Lauren Smolar denied a union link, citing contact volume, wait times, and the liability of non-professional volunteers handling crisis contacts. The four-day gap is documented fact; the motive is contested, the workers and the labor press assert retaliation, the organization denies it, and the unfair-labor-practice charges the union filed have no publicly documented outcome, so no finding of retaliation is asserted anywhere in this file.

The service the helpline delivered is the scale worth holding onto. It had run for more than twenty years and fielded nearly 70,000 contacts in 2022 across calls, chats, and texts, with six paid staff, about two supervisors, and up to roughly 200 rotating trained volunteers doing referral and de-escalation. The organization also reported that volume had grown more than 100% over pre-pandemic levels, that it could not respond immediately to 46% of initial contacts, and that message responses lagged six to eleven days. That contact-volume figure and those queue figures are the organization's own, relayed by journalists, and they were part of the case it made for the change, so they should be read as a justification narrative rather than as an independent measurement.

The harm discovery ran entirely outside the institution. The earliest documented warning came in October 2022, when Monika Ostroff, executive director of the Multi-Service Eating Disorders Association, reported problematic Tessa responses. NPR reports that the specific "healthy snack" language she flagged was quickly removed after she reported it. What is absent from the record is any systemic review, output monitoring, or reassessment of the deployment following that report; and the vendor's chief executive said the flagged language had been part of Tessa's pre-scripted content rather than generated, which the research team denies, so its provenance is contested and the report cannot be read as an early alarm about the generative layer. Then, in late May 2023, fat activist Sharon Maxwell and psychologist Alexis Conason separately tested Tessa and published screenshots. It had recommended a 500 to 1,000 calorie daily deficit, a loss of one to two pounds a week, a 2,000 calorie cap, regular weigh-ins, and where to buy skinfold calipers to measure body composition. Maxwell's account was direct: "Every single thing Tessa suggested were things that led to the development of my eating disorder."

The institution's first response was denial. Communications vice president Sarah Chase commented "this is a flat out lie" on Maxwell's Instagram post and deleted the comment after Maxwell sent screenshots. Early messaging also relayed the vendor chief executive's claim of a 600% traffic surge and "nefarious activity from bad actors trying to trick Tessa." The vendor's own account of the outputs was mixed rather than concessive: Rauws said "We are still trying to determine how a closed system allowed this type of content to be delivered." So the attribution of this advice to the generative layer is what the nonprofit and the research team report and attribute, not an agreed or adjudicated finding about how the words were produced. Five days after the first viral post, on May 30, 2023, the organization suspended Tessa indefinitely, two days before it was to become the sole channel. The helpline closed on June 1 as scheduled anyway. For a period, neither the human service nor its replacement existed, and people arriving were directed to external referrals. As of a December 2023 retrospective, Tessa had not returned and the helpline had not been reinstated.

Two documentary boundaries apply throughout. No regulator, inspector general, or court record exists for the safety failure: the tool was positioned as a wellness and prevention resource, outside device oversight, and the documentary base is concordant independent journalism plus a peer-reviewed trial. And the harmful outputs were documented by adversarial testers and a peer-organization director, not from routine-user logs, so the prevalence of harmful responses in ordinary use is unknown.

The sociotechnical reading

The seam this case is about is not the model's accuracy. It is that the artifact and its evidence came apart, and nothing in the system was watching the gap.

Read structurally, the deployment had two layers with different evidence status. One was a closed, pre-scripted prevention program with a published randomized trial behind it. The other was a generative answering feature added afterward by the vendor, which the trial never covered. Those two layers shipped inside one product, under one name, on one organization's public site, and the trial kept traveling with the name. That is the whole mechanism: a validated warrant stayed attached to an artifact it no longer described. Nothing had to fail dramatically for that to be dangerous. The advice that eventually surfaced - calorie deficits, a weekly rate of weight loss, calipers - is unremarkable consumer dieting content, which is why the reported attribution to a generative layer is unsurprising, and exactly why it is contraindicated for the population that was reading it.

The second structural fact is where discretion sat. In a caseworker system, discretion is highest at the point of service: a human sees each output and can decline it. Here it was inverted. At runtime there was no human who could see or countermand an individual reply, and no real-time monitor for the fifteen months the tool ran. The only available intervention was binary suspension of the whole system, which is what eventually happened. Meanwhile the vendor held effective discretion over what the system could say, and the client's claimed approval gate did not bind it. So the authority to change the product's behavior sat with the party that did not carry the reputational or clinical exposure, and the party that did carry it had one switch, no dial, and no instrument reading.

The third fact is about how the institution learned. Every documented warning arrived from outside: a peer organization's director testing the tool in October 2022, then a psychologist and a community member testing it in May 2023 and posting screenshots. The October report did produce something - the specific language flagged was removed - and that narrowness is the finding. A patch answers the sentence someone quoted; it does not ask what the report implies about everything nobody has quoted yet. Between the private report and the system-level action sat roughly seven months, and what closed the gap was publicity, not evidence age. The institution's own first move was to contradict the report before testing it, which cost it the five days between the first viral post and the suspension.

Then the two clocks. The incumbent service was ended by a labor and cost decision; the replacement was ended by a safety decision; neither decision was aware of the other's timing. Both landed, and served capacity for roughly 70,000 annual contacts fell to referral links. That is worth naming as a governance failure mode in its own right: when a replacement and an incumbent are retired on independent clocks, a system can arrive at a state neither decision intended and no one owns.

The governing levers this case points at follow from all four facts and not from the model. Make the authority to change what the system can say explicit and contractual rather than contested after the fact. Label the published warrant with the population and purpose it actually covers, so a real trial cannot travel further than it went. Evaluate the configuration that shipped and not only the one that was tested. Build a route that turns an external report into a systemic question rather than a content patch, and a blame-free channel so the first institutional response is to test the claim rather than deny it. Keep a human reading the output, because otherwise the first time anyone inside learns what the system says is when someone outside posts a screenshot. And hold the honest boundary throughout: the people who typed into that widget or called that line are not modeled anywhere here, no clinical outcome is derived from any diagram, and the differential risk that ordinary dieting advice carries for people with eating disorders is a documented external fact, never a computed one.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

Independently catalogued as AI Incident Database Incident #545

nprshots2023GroundingInvestigativeSave

NPR, An eating disorders chatbot offered dieting advice, raising fears about AI in health (2023) https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

Appears in: PAN framework development

Grounds: domain grounding: mental and behavioral health chatbots; model org: neda_tessa_chatbot

fitzsimmonscraft2022GroundingAcademicSave

Fitzsimmons-Craft, Chan, Smith, Firebaugh, Fowler, Topooco, DePietro, Wilfley, Taylor, Jacobson, Effectiveness of a chatbot for eating disorders prevention: A randomized clinical trial (International Journal of Eating Disorders, 2022;55(3):343-353; PMID 35274362) https://pubmed.ncbi.nlm.nih.gov/35274362/

https://pubmed.ncbi.nlm.nih.gov/35274362/

Grounds: model org: neda_tessa_chatbot

eliot2023GroundingTrade pressSave

Eliot, Key Lessons Involving Generative AI Mental Health Apps Via That Eating Disorders Chatbot Tessa Which Went Off The Rails And Was Abruptly Shutdown (Forbes, 2023) https://www.forbes.com/sites/lanceeliot/2023/12/13/key-lessons-involving-generative-ai-mental-health-apps-via-that-eating-disorders-chatbot-tessa-which-went-off-the-rails-and-was-abruptly-shutdown/

https://www.forbes.com/sites/lanceeliot/2023/12/13/key-lessons-involving-generative-ai-mental-health-apps-via-that-eating-disorders-chatbot-tessa-which-went-off-the-rails-and-was-abruptly-shutdown/

Grounds: model org: neda_tessa_chatbot

aiincidentdatabase2023GroundingReferenceSave

AI Incident Database, Incident 545: Chatbot Tessa gives unauthorized diet advice to users seeking help for eating disorders (2023) https://incidentdatabase.ai/cite/545/

https://incidentdatabase.ai/cite/545/

Grounds: model org: neda_tessa_chatbot

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalFour days after a labor board certified its paid helpline staff's union election, a national eating-disorder n…

Four days after a labor board certified its paid helpline staff's union election, a national eating-disorder nonprofit told those staff they were terminated and that a chatbot would replace a helpline that had fielded nearly 70,000 contacts in 2022 with six paid staff, about two supervisors, and up to roughly 200 trained volunteers; the chatbot was suspended on May 30, 2023, two days before it was to become the sole channel, after testers published screenshots of it recommending a 500 to 1,000 calorie daily deficit, a 1 to 2 pound weekly loss, a 2,000 calorie cap, regular weigh-ins, and where to buy skinfold calipers, and the helpline closed on June 1 as scheduled while the chatbot was already offline.

nprshots2023GroundingInvestigativeSave

NPR, An eating disorders chatbot offered dieting advice, raising fears about AI in health (2023) https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

Appears in: PAN framework development

Grounds: domain grounding: mental and behavioral health chatbots; model org: neda_tessa_chatbot

EmpiricalThe deployed chatbot had two layers with different evidence status: a closed, pre-scripted program the vendor …

The deployed chatbot had two layers with different evidence status: a closed, pre-scripted program the vendor and the research team describe as unable to depart from its authored content, and a generative question-and-answer feature the operating vendor added in what its chief executive called a systems upgrade covered by the client's contract, a reading the client's chief executive denies by saying the organization was never advised of the changes and would not have approved them; the vendor's public account of the harmful outputs was mixed, saying it was still trying to determine how a closed system allowed such content, so the attribution of the May 2023 advice to the generative layer is the reported and attributed explanation rather than an established mechanism, and the contract scope stays contested with no adjudicated breach.

nprshots2023GroundingInvestigativeSave

NPR, An eating disorders chatbot offered dieting advice, raising fears about AI in health (2023) https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

Appears in: PAN framework development

Grounds: domain grounding: mental and behavioral health chatbots; model org: neda_tessa_chatbot

EmpiricalThe earliest documented external warning about harmful chatbot responses came in October 2022 from the executi…

The earliest documented external warning about harmful chatbot responses came in October 2022 from the executive director of a peer eating-disorder organization, and the specific language she flagged was quickly removed after she reported it, with no documented systemic review, output monitoring, or reassessment of the deployment following; the vendor's chief executive said the flagged language was part of the pre-scripted content rather than the generative layer, which the research team denies, so its provenance is contested, and system-level action arrived roughly seven months later when public screenshots circulated.

nprshots2023GroundingInvestigativeSave

NPR, An eating disorders chatbot offered dieting advice, raising fears about AI in health (2023) https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

https://www.npr.org/sections/health-shots/2023/06/08/1180838096/an-eating-disorders-chatbot-offered-dieting-advice-raising-fears-about-ai-in-hea

Appears in: PAN framework development

Grounds: domain grounding: mental and behavioral health chatbots; model org: neda_tessa_chatbot