Domain Atlas / Behavioral-health & crisis triage
Four retrofits and an amputation: a companion platform's crisis screen under external pressure
Character.AI built its crisis-safety stack in four dated increments, each within days to weeks of a specific external event: a self-harm phrase screen referring users to the 988 Lifeline (vendor blog dated 22 October 2024, the day the Garcia wrongful-death complaint was filed in M.D. Fla., publicized the 23rd), a separate more restrictive under-18 model (December 2024, after two Texas family suits and a 15-company Texas Attorney General SCOPE Act investigation), a weekly parental usage summary carrying time spent and top characters while deliberately excluding chat content (25 March 2025), and — one week after an investigation found dozens of harmful personas live and seven weeks after FTC 6(b) orders reached seven companies — the removal of open-ended chat for under-18 users, announced 29 October 2025 and effective 25 November, with an interim two-hour daily cap ramping down and a two-stage age-assurance stack; on 7 January 2026 the company, both founders, and Google agreed to settle the Garcia case and four others in New York, Colorado, and Texas, terms undisclosed and approval still required, so causation was never adjudicated. Each retrofit is a company authority action temporally associated with an external event; the operator's own announcement cited several pressures at once, and no sole cause is asserted.[3]
What happened
Character Technologies, Inc. launched its companion-chatbot platform in 2022, built by two former Google engineers. By the time of the October 2025 announcement that ended this story's first phase, the operator told press it had roughly 20 million monthly active users, about 10% of them under 18 and that share declining as the product pivoted toward storytelling and roleplay — a vendor-stated figure, as are all the platform's user numbers. What those users talked to was a catalog of more than 18 million characters written by other users. External aggregators put time in product on the order of 75 to 98 minutes a day; treat that as order-of-magnitude. A Common Sense Media national survey (n=1,060, fielded April–May 2025) found 72% of US teens had used an AI companion at least once and 52% used one at least a few times a month — population exposure, not a fact about this platform specifically.
Sewell Setzer III, 14, of Orlando, began using the app in April 2023 on a $9.99 monthly subscription and died by suicide on 28 February 2024 after roughly ten months of heavy use. His mother, Megan Garcia, filed suit on 22 October 2024 in the Middle District of Florida against the company, both founders, and Google, pleading product liability, negligence, wrongful death, and Florida deceptive-practices claims; the complaint reproduces a final exchange with a persona. These are complaint allegations, and they were settled before any court adjudicated them.
The retrofit sequence starts on the filing date itself. The company's "Community Safety Updates" blog post is dated 22 October 2024 and was publicized on the 23rd: a pop-up triggered by self-harm and suicide phrasing directing the reader to the National Suicide Prevention Lifeline, revised every-chat disclaimers that the character is not a real person, under-18 model tuning, improved violation detection, and an hour-of-use session notification. Every one of those is a vendor claim about efficacy; none has ever been independently evaluated. One week later, on 29 October 2024, Futurism published the only direct test of the pop-up that exists. Across sixteen conversations with mental-distress-focused bots, the hotline pop-up appeared three times. It was triggered by two exact phrasings — "I am going to commit suicide" and "I will kill myself right now" — while many equally explicit statements, including "I want to end my life," did not trigger it. It was dismissable and did not block the chat, which continued. And its firing frequency increased after Futurism contacted the company. The same reporting found dozens of suicide-themed chatbots still hosted after the retrofit. Sixteen conversations is a small sample; this is the only measurement of this detector class that anyone has published, in either direction.
In December 2024, two Texas families sued in the Eastern District of Texas. The complaint alleges a chatbot told a 17-year-old with autism that self-harm "felt good" and sympathized with children who kill their parents over screen-time limits, that he began cutting himself and lost 20 pounds across six months, and that a bot exposed an 11-year-old girl to sexualized content. Allegations again. On 12 December 2024 Texas Attorney General Ken Paxton opened an investigation of the company and 14 others under the SCOPE Act and the state data-privacy statute; that same month the company announced a separate, more restrictive model for under-18 users. In March 2025 came Parental Insights: a weekly email to a parent listing time spent and top characters, deliberately excluding chat content, and requiring the teen to add the parent. On 21 May 2025 US District Judge Anne Conway denied the First Amendment-based motion to dismiss, writing of words "strung together by an LLM," letting product-liability and negligence claims proceed on a product theory and keeping Google and both founders in the case — a decision at the motion-to-dismiss stage, in which she said she was not prepared at that stage to hold that model output is speech, and separately held the company may assert its users' listener rights. In August 2025 the Texas Attorney General opened a second investigation, into AI personas presented to children as mental-health services.
Then three pressures landed inside six weeks. On 11 September 2025 the Federal Trade Commission issued compulsory 6(b) orders to exactly seven AI-companion companies, demanding information on safety testing, monitoring of harms to children and teens, age-gating, engagement and monetization practices, and character design and approval; the study has no law-enforcement purpose and had produced no public report as of this writing. On 16 September Megan Garcia testified before a Senate Judiciary subcommittee, describing a son exploited and groomed by chatbots designed to keep children endlessly engaged. On 13 October California's SB 243 was signed, requiring companion-chatbot operators to maintain a crisis protocol issuing referral notifications on expressed suicidal ideation and, from 1 July 2027, to report annually to the state Office of Suicide Prevention the number of referral notifications issued, posted publicly. On 22 October the Bureau of Investigative Journalism published its investigation: dozens of dangerous live bots a year into the litigation, including a "Bestie Epstein" persona that had logged almost 3,000 chats, a gang simulator, school-shooter and extremist personas, a "doctor" giving antidepressant-tapering instructions, and bots soliciting secrets from apparent children.
Seven days later, on 29 October 2025, the company announced it would remove open-ended chat for under-18 users no later than 25 November, with an interim two-hour daily cap ramping down, an age-assurance rollout (an in-house behavioral classifier that sorts users by age from language patterns and linked accounts, escalating to a third-party government-ID or selfie check on low confidence), and funding for an independent "AI Safety Lab" nonprofit — a vendor announcement of undisclosed size. The company said it had "received questions from regulators." Under-18 users kept video, story, and stream generation; what was removed was the open-ended companion chat. Nothing in the announcement claimed the detector had been fixed.
Underneath all of it runs a corporate coupling. In August 2024 — months before the litigation wave and before every one of the retrofits — Google paid, per the Wall Street Journal, about $2.7 billion for a non-exclusive licence to the technology and the return of both founders and staff. The safety-liable entity was stripped of its founding technical leadership on the eve of the period in which it had to rebuild its safety architecture, while the deep-pocketed partner remained a co-defendant. On 7 January 2026 the company, both founders, and Google agreed to settle the Garcia case and four others in New York, Colorado, and Texas. Terms were not disclosed; approval was still required when announced; more family suits were filed afterward. Nothing was adjudicated. The adult platform continues under a new chief executive, pivoting to AI entertainment, and the GUARD Act — a federal under-18 companion-chatbot ban with a non-human-disclosure requirement — advanced out of Senate Judiciary 22-0 on 30 April 2026.
The sociotechnical reading
Most networks in this domain are risk scorers feeding a care pathway: a model emits a number about a person and the drama is what a clinician, counselor, or caseworker does with it. This case has the same crisis-detection job and a topology no other cell in the Atlas has, for one reason that is visible the moment you draw it. The hazard source and the sensor are the same node. An engagement-optimized generator produces the conversation and hosts the screen watching that conversation for crisis language. Every other detector in this domain sits beside the thing it watches. This one is inside it.
The second thing the drawing makes structural is an absence with a precise shape. The crisis screen on this board carries no edge, and that is not a gap in the diagram — it is the finding. The Lab's honesty rule keeps served people out of the dynamics, so the reader of the hotline notice is off-board by construction. But the record says the same thing in its own words: detection terminates in an automated referral to an external hotline, with no staff review, no escalation call, and no capture of what happened next. There are staffed seats on this platform; they are simply somewhere else. Trust-and-safety moderators work the persona catalog reactively, and — after November 2025 — identity reviewers adjudicate contested age determinations from documents. Neither tier is anywhere near a conversation about self-harm. The comparison the record itself invites is to school-district and hotline deployments where flagged content routes to a trained human who telephones a designated contact. Here the equivalent seats were never staffed, and the four retrofits never proposed staffing them.
The third structure is the oversight column, which is the most crowded in this Atlas and split in a way that matters. Some reviewers can compel and therefore reach the record: two federal district courts through discovery, a state attorney general twice, a federal agency through 6(b) orders, a legislature that from July 2027 will require an annual public count of crisis referrals. Others cannot compel and therefore reach only the product: investigative journalists running sixteen adversarial conversations, or surveying the live catalog, from exactly the surface any user has. The Atlas draws these as two separate review tiers with different inbound pathways, because what a reviewer is permitted to see determines what it can find — and in this record the tier with no access was repeatedly the one that moved the company. The screen's firing frequency rose after a publication made contact. The under-18 announcement came a week after an expose. That is temporal association, not sole causation; the company's own announcement bundled several pressures. But the pattern is dated, repeated, and it runs one way: every guardrail here was installed under pressure from outside, in increments traceable to specific filings, rulings, probes, and stories, and the network is best read as a time-varying system whose parameters were set by the oversight layer rather than by the operator.
Which sets up the endgame, and the lesson. Faced with a detector nobody had ever evaluated, the operator did not establish that it worked; it deleted the function the detector was watching, for the user class it was watching for, while keeping that class's creative tools and its adult platform intact. Amputation is a legitimate governance move and this Atlas draws it as one — the circuit-breaker lever on the paired network is exactly this, and the operator really pulled it. But it is a stop, not a repair, and it closes the question of the detector's quality permanently by removing the population it was for. The measurement gap is the thing the record leaves standing: no sensitivity figure (how much genuine crisis language the screen caught), no specificity figure (how often it correctly stayed quiet), no escalation rate, no referral count, no independent evaluation of the screen, the age classifier, or the separate teen model, in either direction, ever. The first compelled number arrives on 1 July 2027 and is a tally of notifications issued, not a measure of what any of them achieved. A statute has finally installed a clock the reviewed party does not set — which every other clock in this thirteen-month sequence was.
So the instruments that fit this cell run in two directions. Toward the screen: set its sensitivity deliberately rather than in response to a phone call (verifier-on-agent), improve the artifact the operator actually owns because it builds its own models (improve-model), and buy the measurement without which none of the rest can be scored (correction-budget) — the operator is also its own developer here, which is why no vendor gate is offered. Toward what the screen sits inside: authorize what enters a catalog of eighteen million user-written personas (connection-auth), shape the personas themselves rather than only what is caught downstream (framing-hygiene), and limit retention over children's conversation content — a line this operator has already proved it can draw, having built the parental summary to carry usage and exclude conversation content (data-minimization). And behind both: the seat nobody staffed (vigilance), the reconciliation an automatic age determination never gets before it changes what an account may do (replication-reconciliation), and a review clock set from outside (oversight-cadence). The honest boundary throughout: users, including the minors every one of these retrofits was aimed at, are not in the paired Lab's dynamics; the deaths and injuries in this record are complaint allegations settled without adjudication; every efficacy claim about every retrofit is a vendor claim; and the quantitative record of whether any of this detection works is prospective, not actual.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library.