PAN Lab example
Character.AI crisis-safety stack
The screen with nobody behind it
A companion-chat platform serving roughly twenty million people, about a tenth of them under eighteen, built its entire crisis-safety layer while under external pressure, in four dated steps. A phrase screen raising a hotline notice arrived the same day the first wrongful-death complaint was filed. A separate, more restrictive model for teenagers followed two family suits and a state investigation. A weekly parental summary followed in the spring. Then, one week after an investigation found dozens of harmful personas live on the platform and seven weeks after a federal agency compelled seven companies to answer for how they protect minors, the operator removed open-ended chat for under-eighteen accounts altogether. Read the board and the shape of that history becomes structural. The generator both produces the conversation and hosts the screen watching it. The one measurement anyone has of that screen found it firing three times across sixteen adversarial conversations, on two exact phrasings, dismissable and without blocking the chat - and firing more often after the testers made contact. And behind the screen there is nobody: the referral goes to an external hotline, and no seat anywhere on this board is staffed to follow it. The people who do work here moderate a catalog of eighteen million user-written characters after they are reported, and check identity documents when someone contests an age. Five suits settled in January 2026 without any adjudication of causation. Before you pick a target level: this board cannot be won under All Governance Targets. Cost is not what blocks it. Lift the pathway requirement on its own and the All Governance Targets level can be met from a stack costing 7 of the 13 you have, with the benefit reading above the margin it must clear. The pathway gate is the only gate that fails. The Service and Safety Targets level closes every pathway at exactly the full budget. The All Governance Targets level adds path dependence, and what reopens under it is the generator writing into its own stores: the retained conversation log, and the account file it reads ages from. No legal configuration closes those and clears the rest at the same time. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore, Service Targets Only, and Service and Safety Targets can be won, and the Service and Safety Targets level is won at exactly the full budget.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Companion-platform-class crisis screen under external review network: 12 components and 25 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 1 assumed · 12 published baseline · 2 measured. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- baseline
This models the litigation-period companion-platform pattern documented in the Character.AI case file. The four safety increments are company authority actions, each dated within days to weeks of a specific filing, ruling, state investigation, federal compulsory order, or published investigation. That is temporal association: the operator's own under-18 announcement cited several pressures at once, and nothing here asserts sole causation for any of them.
- measured
The crisis screen is drawn and described from the one direct measurement that exists, and no further. Adversarial testing published in October 2024 found the hotline notice appeared 3 times across 16 conversations with distress-focused bots, triggered by two exact phrasings while many equally explicit statements did not trigger it, dismissable and non-blocking so the same conversation continued, with firing frequency rising after the publication contacted the company. n is 16, so this is anecdotal rather than a measured rate, and it is the only measurement of this detector in either direction.
- baseline
The crisis screen carries no edge, and that is the derivation rather than a gap in it. This Lab's honesty invariant puts served people outside the dynamics: U is the operator network. The people the crisis notice is shown to are users, including minors, so there is no operator for the screen to check output to. The record says the same thing in its own words: detection terminates in an automated referral to an external hotline with no staff review, no escalation call, and no capture of what happened next.
- assumed
Users are absent from this diagram by construction, which removes a documented dynamic from the board rather than denying it. The engagement-optimized agreement pattern at the centre of the complaints runs between the generator and people who are not in the dynamics, so it is narrated in the case file and never computed here. For the same reason no operator-to-model framing pressure is drawn: the only operator-to-model pathway on this board is the moderation tier's takedowns and tuning.
- baseline
The catalog inflow is drawn at the top of the range because the personas are what the generator plays rather than a reference corpus it consults: more than 18 million of them, user-authored, with entry ungated and moderation documented as reactive. The October 2025 investigation is the evidence that what sits in this catalog reaches users at strength, having found dozens of harmful personas live a year into the litigation, one of which had logged almost 3,000 chats.
- baseline
The machine write into the conversation record is drawn at the top of the range on a structural fact, not an analogy: generation and recording are a single act, with no draft-then-approve step anywhere in the pathway. That is precisely why complete transcripts existed to become the evidentiary core of the lead complaint. No retention limit and no reconciliation process for that store appears anywhere in the record in either direction.
- baseline
The age determination is drawn on the enforcement idiom because the record documents an automatic account action: an interim two-hour daily cap announced 29 October 2025, ramping to removal of open-ended chat for under-18 accounts on 25 November 2025, applied from the determination with no human step in between. The reconciliation pathway back into the record is drawn empty because the documented remedy is a contest with identity documents that runs after the change has already applied.
- baseline
The oversight tier is split into two review steps on a documented difference in access, not to create one. Courts, agencies and legislatures reach the record through discovery and compulsory information demands, so they are wired in on a store-side inbound at a substantial level. Outside investigators had the public product surface and nothing more, so they are wired in on the model-side inbound at a substantial level, which is what repeated adversarial use of a live consumer product is worth. Neither tier ever operated inside the loop.
- baseline
The statutory reporting pathway is drawn empty because it has not begun. A California statute signed 13 October 2025 requires a crisis protocol issuing referral notifications and, from 1 July 2027, an annual report of the number of notifications issued to a state office, posted publicly. That is the first compelled quantitative record of this detector class, and it is a count rather than a measure of what the referrals achieved.
- baseline
The model-side check runs empty because the record documents no independent evaluation of the crisis screen, the age classifier, or the separate under-18 model, in either direction. Every efficacy claim about the four retrofits is a vendor claim. The nearest substitutes are investigative journalism and an outside risk assessment, neither of which had access to the system, and both of which are drawn as the second reviewer rather than as a model check.
- measured
The monoculture self-loop carries a specific measured shape rather than a general concern: one screen sits across every conversation on the platform, so a phrasing outside its list is a gap in all of them at once. That is what the adversarial testing observed when two exact phrasings triggered the notice and many equally explicit statements did not. The separate under-18 model is the same builder's variant of the same stack, and no evaluation compares them.
- baseline
Both data-leaving pathways are deliberate product features rather than leaks, and both are drawn faint because both are conditional. Identity documents reach a third-party provider only on a low-confidence determination, and the weekly summary reaches a parent only where the teen has added them. The summary's exclusion of conversation content is the single documented data-minimisation control anywhere in this record, which is why the diagram separates the account record from the conversation record rather than drawing one store.
- baseline
A heavy workload against very limited capacity. Demand from roughly 20 million monthly active users, about 10 percent of them under 18 and declining, on a catalog of more than 18 million personas, with third-party aggregations putting time in product on the order of 75 to 98 minutes a day; the user-scale figures are vendor-stated to press and the engagement figures are third-party aggregations, so both are order-of-magnitude. Capacity is very limited because the crisis pathway has no staffed seat at any point: the two operator tiers on this board work content moderation and identity documents, and the dossier's own comparison is to school-district deployments where flagged content routes to trained reviewers who telephone a named contact.
- baseline
No outcome for any user is modeled, and none may be inferred. The deaths and injuries in this record are the subject of complaint allegations that were settled on 7 January 2026 without any adjudication of causation, with approval still required when announced and terms sealed. The population-exposure survey figure, that 72 percent of United States teens have used AI companions and 52 percent use them at least a few times a month, is boundary context recorded in the case file, never a parameter here.
- baseline
The preliminary court ruling is carried as what it was. The order of 21 May 2025 was at the motion-to-dismiss stage: it let product-liability and negligence claims proceed on a product theory, said the court was not prepared at that stage to hold that model output is speech, and kept the founders and the licensing partner in the case. Nothing here presents it as a final holding about protected speech.
What this example does not show
- No outcome for any user is modeled, and none may be inferred from this diagram. The Lab reads institutional propagation only; the people using this platform, including the minors the retrofits were aimed at, are boundary-only. The deaths and injuries in this record are the subject of complaint allegations that were settled on 7 January 2026 without any adjudication of causation, with approval still required when announced and terms sealed.
- The four safety increments are company authority actions, each dated within days to weeks of a specific filing, ruling, state investigation, federal compulsory order, or published investigation. That is temporal association and nothing more: the operator's own under-18 announcement cited several pressures at once, and no sole cause is asserted for any step, including the one that followed an investigation by a week.
- The crisis screen is described here from the only direct measurement that exists and no further: three firings across sixteen adversarial conversations with distress-focused bots, triggered by two exact phrasings while many equally explicit statements did not trigger it, dismissable and non-blocking, with firing frequency rising after the publication contacted the company. That test had a sample of sixteen, so it is anecdotal rather than a rate. No independent evaluation of the screen, the age classifier, or the separate under-18 model exists in either direction, and every efficacy claim about the retrofits is a vendor claim.
- The scale figures are vendor-stated and the engagement figures are third-party aggregations, both used as order-of-magnitude only: roughly twenty million monthly active users with about ten percent under eighteen and declining, told to press by the operator, and time in product on the order of seventy-five to ninety-eight minutes a day from external aggregators. The count of almost three thousand logged chats on one persona is the investigator's own figure and is not rounded up.
- The May 2025 court ruling is carried as what it was: a preliminary decision at the motion-to-dismiss stage that let product-liability and negligence claims proceed on a product theory, said the court was not prepared at that stage to hold that model output is speech, and kept the founders and the licensing partner in the case. It is not a final holding that chatbot output is unprotected.
- The engagement-optimized agreement pattern at the centre of these complaints does not appear on this board. Its two ends are the generator and the user, and users are outside the dynamics by the same honesty invariant that keeps every served population off every Lab diagram. It is narrated in the case file and is never computed here.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Character.AI built its crisis-safety stack in four dated increments, each within days to weeks of a specific external event: a self-harm phrase screen referring users to the 988 Lifeline (vendor blog dated 22 October 2024, the day the Garcia wrongful-death complaint was filed in M.D. Fla., publicized the 23rd), a separate more restrictive under-18 model (December 2024, after two Texas family suits and a 15-company Texas Attorney General SCOPE Act investigation), a weekly parental usage summary carrying time spent and top characters while deliberately excluding chat content (25 March 2025), and — one week after an investigation found dozens of harmful personas live and seven weeks after FTC 6(b) orders reached seven companies — the removal of open-ended chat for under-18 users, announced 29 October 2025 and effective 25 November, with an interim two-hour daily cap ramping down and a two-stage age-assurance stack; on 7 January 2026 the company, both founders, and Google agreed to settle the Garcia case and four others in New York, Colorado, and Texas, terms undisclosed and approval still required, so causation was never adjudicated. Each retrofit is a company authority action temporally associated with an external event; the operator's own announcement cited several pressures at once, and no sole cause is asserted.
empirical- Vendor Character.AI, Community Safety Updates (vendor blog) (2024) https://blog.character.ai/community-safety-updates/
- Vendor Character.AI, Taking Bold Steps to Keep Teen Users Safe on Character.AI (vendor blog) (2025) https://blog.character.ai/u18-chat-announcement/
- Investigative CNN Business, Character.AI and Google agree to settle lawsuits over teen mental health harms and suicides (2026) https://www.cnn.com/2026/01/07/business/character-ai-google-settle-teen-suicide-lawsuit
The only direct measurement of Character.AI's self-harm detector is a single adversarial test published 29 October 2024: across sixteen conversations with mental-distress-focused bots the 988-hotline pop-up appeared three times, triggered by two exact phrasings ('I am going to commit suicide' and 'I will kill myself right now') while many equally explicit statements including 'I want to end my life' did not trigger it, dismissable and non-blocking so the chat continued, with firing frequency increasing after the publication contacted the company — a sample of sixteen, and therefore anecdotal rather than a measured rate. No independent evaluation of the sensitivity or specificity of the crisis classifier, the two-stage age-assurance classifier, or the separate under-18 model exists in either direction; every efficacy claim about the retrofits is a vendor claim. The first compelled quantitative record of this detector class is prospective: California SB 243, signed 13 October 2025 and operative 1 January 2026, requires a crisis protocol issuing referral notifications and, from 1 July 2027, an annual report to the state Office of Suicide Prevention of the number of notifications issued, posted publicly — a count of referrals, not a measure of what they achieved.
empirical- Investigative Futurism, After Teen's Suicide, Character.AI Is Still Hosting Dozens of Suicide-Themed Chatbots (2024) https://futurism.com/suicide-chatbots-character-ai
- Government California Legislative Information, Senate Bill No. 243, Companion chatbots (2025) https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB243
Character.AI's crisis pathway terminates in an automated referral to an external hotline with no staff review, no escalation call, and no capture of what followed; the human roles the record does document sit elsewhere — trust-and-safety moderators taking down user-authored characters reactively, parents receiving a weekly usage summary with no transcript access and only where the teen adds them, and, after November 2025, third-party identity reviewers adjudicating contested age determinations. The external review tier is by contrast unusually crowded and repeatedly behavior-forcing: two federal district courts, the Texas Attorney General twice, FTC 6(b) compulsory orders to seven companies on 11 September 2025, California SB 243, and a Senate subcommittee hearing, alongside investigative journalism that on 22 October 2025 — a year into the litigation — found dozens of harmful bots live including a 'Bestie Epstein' persona that had logged almost 3,000 chats, a gang simulator, school-shooter personas, and a 'doctor' giving antidepressant-tapering instructions. None of these reviewers ever operated inside the decision loop; what they moved was the structure itself.
empirical- Investigative The Bureau of Investigative Journalism, Gang leaders, school shooters and Bestie Epstein: meet Character.AI's chatbot companions (2025) https://www.thebureauinvestigates.com/stories/2025-10-22/gang-leaders-school-shooters-and-bestie-epstein-meet-character.ais-chatbot-companions
- Government Federal Trade Commission, FTC Launches Inquiry into AI Chatbots Acting as Companions (2025) https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions
- Vendor Character.AI, Introducing Parental Insights: Enhanced Safety for Teens (vendor blog) (2025) https://blog.character.ai/introducing-parental-insights-enhanced-safety-for-teens/
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Verify output — Put a verifier on the agent
- Upgrade model — Improve the model
- Review the riskiest first — Risk-tiered oversight
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Pause AI on alarms — Deployment circuit-breaker
- Escalate checks — State-feedback vigilance
- Vet connections — Connection authorization
- Store less data — Data minimization
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
Documented case histories
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Kaiser Permanente Suicide-Risk Model
- Crisis Text Line & Loris.ai
- LyssnCrisis counselor QA at ProtoCall Services (988)
- NarxCare
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- Tessa chatbot replacing the NEDA eating-disorder helpline
- Woebot (a governed app wind-down)