PAN Lab example
NYC Teenspace
The surface everyone could audit, and the one nobody did
A city bought population access to a commercial teletherapy platform for every teenager who lives there: 26 million dollars over three years, no insurance, no referral, parental consent and a licensed therapist. Inside the therapy chat the vendor runs a language model that reads teen-authored messages for wording consistent with suicide risk and alerts the clinician already treating that teen. It never acts on its own, and the counted rate is small: about 50 flags among about 6,800 users in the first six months, alongside 36 high-risk events. Then look at what could be inspected from outside. Nobody has ever examined the model; the evaluation the city said it was planning in September 2024 had not appeared by mid-2026. What outsiders could read was the public sign-up pages, and a parent coalition read them with a free tracker tool and counted 15 advertising trackers and 34 cookies carrying teen visitors' identifiers to named companies, several of them companies the same city was suing over teen mental health. The vendor stripped the tags on December 11, 2024, about three months after the complaint and a week before the city first said the contract rider would be amended. Within weeks advocates documented tags again on the pages beyond the landing page, including the page holding the revised privacy policy. The audited surface and the algorithmic surface never met. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. Cost is not what blocks it. Inside the budget the best legal settings clear the benefit margins comfortably. Lift the pathway requirement on its own and the board wins at both tiers, from a stack costing 7 of the 10 you have under Service and Safety Targets and 8 under All Governance Targets. The pathway gate is the only gate that fails. What stays open at every affordable price is the therapy record read back by the clinician who wrote it. That pathway is the care itself, and no lever in this deployment's documented authority set closes it. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Teenspace-class purchased-platform teen teletherapy network: 8 components and 19 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 10 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the government-purchaser-of-a-consumer-platform pattern the case file documents, not a reconstruction of the service, and it makes no claim about the scan's accuracy in either direction. Every element is a model of institutional propagation; a safe-looking or unsafe-looking reading here is a property of this model, never a measurement of the deployment.
- baseline
Eight nodes, all documented, none decorative. The record names one model, two operator groups with different authority and different documented outcomes, one record store that is the therapy conversation itself, one external intake-and-web surface, two review tiers reaching different parts of the system, and one off-network sink with measured flow into it. Four kinds are deliberately absent for evidence reasons. No enforcement node: the dossier states the algorithm never acts autonomously and the escalation decision is entirely the clinician's, so no record-driven action system exists to draw. No worklist: no queue, backlog or triage list appears anywhere in the record, and this model directs attention inside one clinical relationship rather than ordering a queue. No guardrail and no automated output screen: nothing checks the model's output before the clinician sees it, and the model's own scan is its output rather than a check of it, so drawing it as a bounded screen would misread which pathway is which. No retriever.
- baseline
The adoption arm is drawn faint, among the lowest in this domain and the only one there set by a count rather than by a reading, and the counted number is this. In the programme's first six months the scan flagged roughly 50 teens as moderate-to-high suicide risk, under one percent of about 6,800 users, while therapists navigated 36 high-risk events. That is a low-volume, high-salience alert, the inverse of the alert-flood shape most behavioural-health networks in this catalogue draw, and it is why this deployment is a poor model of alert fatigue and a good model of something else: an attention signal inside a relationship the clinician already holds. Whether that low rate reflects a conservative threshold or a screen that reads little cannot be settled from the record, and nothing here should be read as settling it.
- baseline
The clinical spine runs at full strength in three directions because in this deployment the record store is the work. The scan reads the message stream as it is written and is the one continuous reader of it; the clinician's reply is kept as sent rather than written up into a separate file; and the clinician reads a teen almost entirely from that written stream. The supporting figures are the vendor's own to investors, labelled as such: more than 90 percent of enrolled teens actively text, more than half use messaging as their only modality, and the platform-wide bank is described as 8 billion words and 140 million message exchanges. The service promise behind them, unlimited asynchronous messaging with replies five days a week and one 30-minute live session a month, is on the department's public programme page.
- baseline
The two data-leaving pathways run from different components and carry different payloads, which is the structural finding of this case rather than a drawing choice. The page's data-leaving pathway leaves the public sign-up surface and carries identifiers: a September 2024 coalition audit counted 15 advertising trackers and 34 cookies with named recipients, the department recorded all such trackers removed on December 11, 2024, advocates documented tags again in January and February 2025 on pages beyond the landing page including the page carrying the revised privacy policy, and a February 2025 technical investigation found the vendor's teen pages for two other cities still transmitting until a reporter made contact. Per the binding roster ruling this pathway is modelled identifier-level and contested-scope: the vendor chief privacy officer's statement that no personal medical information was transmitted and the advocates' dispute of the fix are both recorded, and nothing here asserts that therapy content crossed. It is drawn faint rather than empty because it was counted, and no higher because a dated removal and a clean independent test of the main page sit in the same window as the documented recurrence on secondary pages.
- baseline
The record store's data-leaving pathway is a separate pathway, drawn faint, recorded as open rather than as a measured flow, because none of its three documented mechanisms is evidenced as an event in this teen programme. A court can compel the bank, shown by a 2026 report of an adult platform member's complete message history obtained by her former employer and used in court, which evidences the store's legal exposure and is expressly not a Teenspace incident. The vendor states it uses the same bank to develop behavioural-health language models, a use the purchase contract never addressed. And the bank transfers intact to a new corporate owner under the acquisition announced in March 2026, while the city contract is still running. What the record does not contain is any retention limit for an enrolled teen's messages.
- baseline
The outside review tier is wired to the input surface rather than to the model, and it is drawn stronger than the purchaser's own check, because that is what the record shows. Nobody outside this deployment has ever examined the algorithm. What outsiders could read was the public pages, and they read them repeatedly with a public tracker inspection tool across at least nine months, published dated findings, and were twice followed by a vendor change. The purchaser's check is drawn faint beside it: real, contractual, and used once in the documented period, about three months after the first public complaint, with a follow-up letter that had gone more than a month without an answer. What the dates support is engagement rather than contractual causation, and the shipped case file for this deployment reads them the same way: the trackers came down on December 11, 2024, a week before the city first said the rider would be amended, so the rider stays a forward-looking lever whose effect on conduct nobody has yet measured. An audit function performed by non-governmental organisations and journalists, more actively than by the body that promised one, is the dossier's oversight-substitution finding drawn as structure.
- assumed
The scan's write into the record is drawn faint and this is the one baseline on this board resting entirely on an acknowledged silence in the sources, which is said plainly here rather than dressed as a finding. The vendor describes 6.2 million retained assessments sitting alongside the message bank, so something is retained beside the messages; what the record never says is whether the scan's own flag is among them. A conservative faint value therefore records a real but unquantified write rather than assuming a full write-through, and no inference about what the platform retains about an alert may be drawn from it in either direction.
- baseline
The aggregate counts reaching the purchaser are drawn faint, and that low baseline is the evidence rather than a drawing choice. This is the pathway that wires the purchasing agency into the board at all, and what the record shows travelling along it is vendor-reported totals: about 50 moderate-to-high-risk flags and 36 high-risk events in the first six months. The dossier records alert-outcome feedback to the city as aggregate vendor-reported counts, and documents no per-case visibility for the department and no channel by which it sees an individual alert. A city that bought a suicide-risk algorithm for 26 million dollars learns about it in totals, and the baseline is set to say exactly that much and no more.
- baseline
Four remaining pathways are drawn faint, each a documented channel with a small or twice-counted volume rather than an inferred one. The sign-up details entering the record are faint because what crosses today is age and zip code per the department's own December 18, 2024 letter, a rare case of a documented remediation setting a baseline down. The vendor staff's control of the sign-up fields and page tags is faint because the record shows that channel used twice: the December 11, 2024 removal and minimisation, and further removals on the vendor's teen pages for two other cities after a reporter made contact in February 2025. The vendor's tuning of the scan is faint because the model is proprietary and vendor-built and the only documented change in the period is a further clinical agent the vendor told investors was in beta in February 2026. The clinician escalations into vendor safety handling are faint because the counted volume is 36 high-risk events in six months: real, and small.
- assumed
Three pathways are drawn empty because the record documents the absence, and one of them rests on silence, which is said plainly here rather than dressed as a finding. The independent programme evaluation was promised in a September 12, 2024 email and had not been published as of mid-2026, and no inspector-general, comptroller or federal consumer-protection action specific to this programme was found. The outside test of the model does not exist either: the one performance figure in the record is the vendor's 83 percent agreement with a human expert, taken from a 2020 vendor-affiliated study of adult platform data and never a measurement of teens. The clinician-feedback pathway is the one resting on silence: the sources describe no channel by which the clinicians who receive alerts could correct or retune the model, so a conservative empty value is used, and the adjacent documented fact is that the purchasing department has no reach into the algorithm or the alert flow either.
- baseline
The model self-loop is drawn at a substantial level rather than the catalogue's usual low setting for a scale reason that is documented rather than assumed. This is not one model per agency: one proprietary model reads every member of an entire commercial platform, which the vendor puts at about 32,000 members flagged since 2019, and the same platform runs near-identical teen services in other cities. The one published performance figure for it comes from a 2020 vendor-affiliated study of adult platform data, and no measurement against this teen population exists anywhere in the record. The record does not describe the training corpus beyond anonymised consented platform transcripts, so the adult scoping is stated of the 2020 study only, per the binding roster ruling, and is not asserted of the corpus. A gap in how it reads one kind of wording is therefore the same gap in every conversation at once. The paired absence is drawn beside it: no independent test of the model against this population exists.
- baseline
A heavy workload against limited capacity is derived from counted workload against a documented service cadence, not from the modal pair. Demand: registration ran from about 6,800 at six months to more than 19,000 by December 1, 2024 on the department's own letter, to a vendor-reported 45,000 or more by February 2026, roughly a sixfold rise in 21 months, against a promise of unlimited asynchronous messaging that more than 90 percent of enrolled teens take up. Capacity, and deliberately neither low nor high: the counterfactual is narrow, because the automated part is the risk-language scan and not the therapy, so without it a licensed clinician still reads their own client's messages, which is a real working manual process. It is held below full because replies run five days a week over an unlimited messaging channel, so there are windows in which nobody is reading, and because a school social worker publicly characterised the modality as interim care rather than a substitute for intensive clinical work.
- assumed
Every performance and outcome figure in this record is a vendor or city self-report and is labelled at each use, per the binding roster ruling. The 83 percent detection figure rests on a 2020 vendor-affiliated study of adult platform data and is never a Teenspace measurement. The 65 percent reported improvement figure at six months and the 66 percent measurable clinical improvement figure at 45,000-plus enrollees have undisclosed instruments and no independent verification. The 400,000 to 500,000 eligible-teen figure is a vendor investor claim, not a city figure. The programme term is stated as a three-year programme launched November 15, 2023, because the department's programme page carries no end date. The class action over one tracker's fingerprinting was filed August 20, 2024 and withdrawn without prejudice in September 2025 after the judge signalled an inclination to dismiss, so no holding of any kind exists and none may be implied. Continuation past the contract term and the effect of the acquisition on this contract were not publicly documented as of mid-2026.
- assumed
Served teens are not in the dynamics. The Lab reads institutional propagation only; a message, a flag or an escalation on this map is an institutional signal and never a person, and no clinical, safety or suicide outcome is computed anywhere on this diagram. The record does document reach in unusual detail, and that detail lives in the case file rather than here: about 80 percent of early registrants Black, Latino, Asian American or Native American, about 70 percent female, and more than half resident in the city's priority equity neighbourhoods, with the vendor later reporting 82 percent BIPOC and nearly 45 percent in high health and income disparity areas. Those are access statistics. No differential outcome, differential harm or subgroup performance measurement for this programme exists anywhere in the record, so no demographic parameter enters this model and none may be inferred from it.
What this example does not show
- No outcome for a served teen is modeled. The Lab reads institutional propagation only; the teenagers in this case are boundary-only, and no clinical, safety or suicide outcome for any of them is computed on this diagram. The record does document reach in detail - about 80 percent of early registrants Black, Latino, Asian American or Native American, about 70 percent female, more than half in the city's priority equity neighborhoods, with the vendor later reporting 82 percent BIPOC - and those are access statistics that live in the case file. No differential outcome, differential harm or subgroup performance measurement for this program exists anywhere in the record, so none is shown or implied here.
- Every performance and outcome figure in this record is a vendor or city self-report and is labeled at each use. The 83 percent detection figure rests on a 2020 vendor-affiliated study of adult platform data and is never a measurement of this teen program. The 65 percent reported improvement figure at six months and the 66 percent measurable clinical improvement figure at 45,000-plus enrollees have undisclosed instruments and no independent verification. The 400,000 to 500,000 eligible-teen figure is a vendor investor claim rather than a city one. The independent evaluation the health department said it was planning in September 2024 had not been published as of mid-2026, and that absence is the oversight story rather than a gap in this account.
- The tracker leakage is modeled identifier-level and contested-scope, exactly as the record leaves it. What was measured leaving the public pages was visitor addresses, device and browser identifiers and referral data; the vendor's chief privacy officer stated that no personal medical information was transmitted, and advocates disputed how complete the remediation was. Both positions are carried and neither is resolved here. Nothing on this diagram asserts that therapy content crossed the boundary.
- Nothing was adjudicated. The class action over one tracker's fingerprinting of device, geographic, referral and URL data including minors was filed August 20, 2024 in the Central District of California and withdrawn without prejudice in September 2025 after the judge signaled an inclination to dismiss: there is no holding of any kind and none may be inferred. The 2026 report of a subpoenaed therapy history concerned an adult employer-benefit member of the platform, not a teen in this program; it evidences the record store's legal exposure and is not an incident in this deployment. No inspector-general, comptroller or federal consumer-protection action specific to this program was found, and absence of findings is absence of evidence rather than clearance.
- The program is stated as a three-year program launched November 15, 2023, because the health department's own program page carries no end date. Continuation past the contract term, and what the pending acquisition of the vendor does to the contract or to the message bank the acquisition transfers, were not publicly documented as of mid-2026 and are not modeled here.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
New York City launched NYC Teenspace on November 15, 2023 as a three-year, $26M contract buying population access to Talkspace's existing consumer teletherapy platform for all city residents aged 13-17, with unlimited asynchronous messaging (therapist replies five days per week) plus one 30-minute live session per month from NY-licensed clinicians; inside the therapy chat a proprietary NLP model scans teen-authored messages for suicide and self-harm language and fires an urgent real-time alert to the teen's own treating therapist, which never acts autonomously — escalation (child-protective-services referral, intensive therapy, hospitalization) is entirely the clinician's; in the first six months the scan flagged roughly 50 teens as moderate-to-high suicide risk (under 1% of about 6,800 users) while therapists navigated 36 high-risk events including suicide attempts, a child-abuse report, and a drug overdose; enrollment ran from about 6,800 (May 2024) to 19,000+ (DOHMH letter, Dec 1 2024) to a vendor-reported 45,000+ (Feb 2026), with about 80% of early registrants Black, Hispanic, AAPI, bi-racial, or Native American, about 70% female, and more than half resident in the city's priority TRIE neighborhoods.
empirical- Government NYC Office of the Mayor, Mayor Adams and DOHMH Commissioner Dr. Vasan Launch NYC Teenspace, a Tele-mental Health Service for NYC Teens (2023) https://www.nyc.gov/office-of-the-mayor/news/869-23/mayor-adams-dohmh-commissioner-dr-vasan-launch-teenspace-tele-mental-health-service-nyc
- Investigative Chalkbeat New York, Six months in, NYC free online therapy platform for teens has seen 7,000 signups (2024) https://www.chalkbeat.org/newyork/2024/05/23/thousands-of-teens-sign-up-for-free-online-therapy-talkspace/
- Government New York City Department of Health and Mental Hygiene, Letter to the Parent Coalition for Student Privacy, NYCLU and AI for Families re: Teenspace Program (2024) https://studentprivacymatters.org/wp-content/uploads/2025/01/DOHMH-Teenspace-Letter-12.18.2024-003.pdf
- Government New York City Department of Health and Mental Hygiene, NYC Teenspace program page and FAQ for Parents and Guardians (2026) https://www.nyc.gov/site/doh/health/health-topics/teenspace.page
A September 2024 advocacy tracker audit of the NYC Teenspace pages counted 15 ad trackers and 34 cookies sharing teen visitors' personally identifiable information with recipients including Facebook/Meta, Amazon, Google, and Microsoft, alongside sensitive intake data (name, date of birth, address, school, gender, mental-health screening answers) collected before parental consent was secured; DOHMH's December 18, 2024 letter (General Counsel Landau, Chief Privacy Officer Elcock) recorded that Talkspace removed all social-media and advertising trackers as of December 11, 2024, that the sign-up flow was minimized to age and zip code only, and that the contract's Data Security Rider would be amended to ban marketing use and trackers outright; advocates then documented continuing tracker-based disclosure on pages beyond the landing page (including the page hosting the revised privacy policy) on January 9 and February 12, 2025 with no DOHMH response to their follow-up letter after more than a month, and February 2025 technical testing found the NYC landing page clean as of January 24 while Talkspace's Seattle and Baltimore teen pages still transmitted visitor IP addresses to TikTok, Meta, Snapchat, Google, X, Reddit, LinkedIn, Spotify, and Quora until the reporter's inquiry — several recipients being companies NYC had sued in February 2024 over teen mental-health harm. Talkspace's Chief Privacy Officer stated no personal medical information was transmitted; advocates dispute the completeness of the remediation, so the leakage is identifier-level and contested-scope.
empirical- Investigative Parent Coalition for Student Privacy, Privacy concerns with NYC student use of the Teenspace online counseling service (2024) https://studentprivacymatters.org/privacy-concerns-about-nycs-promotion-of-the-teenspace-online-counseling-service/
- Government New York City Department of Health and Mental Hygiene, Letter to the Parent Coalition for Student Privacy, NYCLU and AI for Families re: Teenspace Program (2024) https://studentprivacymatters.org/wp-content/uploads/2025/01/DOHMH-Teenspace-Letter-12.18.2024-003.pdf
- Investigative Parent Coalition for Student Privacy, Continuing Teenspace privacy violations, despite assurances from city (2025) https://studentprivacymatters.org/continuing-teenspace-privacy-violations-despite-assurances-from-city/
- Investigative Gizmodo, Teen Mental Health App Sent Kids' Data Straight to TikTok (2025) https://gizmodo.com/teen-mental-health-app-sent-kids-data-straight-to-tiktok-2000557615
The only published performance figure for the Teenspace suicide-alert algorithm is Talkspace's own claim of 83% accuracy versus a human expert, resting on a 2020 Psychotherapy Research study of ADULT platform data rather than any measurement of the teen population it reads, and the independent evaluation DOHMH said in a September 12, 2024 email that it was planning had still not been published as of mid-2026, with no OIG, comptroller, or FTC action specific to Teenspace located; outcome figures (65% 'reported improvement' at six months; 66% 'measurable clinical improvement' among 45,000+ enrollees) are vendor or city self-reports with undisclosed instruments and no independent verification; meanwhile the record store the program writes into is described by Talkspace's CEO as 8 billion words, 140 million messages, and 6.2 million assessments, is subpoenable (a Talkspace user's complete therapy history was obtained by her former employer and used in court in 2026 — an adult employer-benefit member, not a Teenspace teen), is stated by the vendor to be used for training behavioral-health LLMs, and transfers intact under the pending $835M Universal Health Services acquisition announced March 9, 2026 and expected to close in Q3 2026, while the NYC contract is still live.
empirical- Vendor Talkspace, Inc. via Business Wire, Proprietary AI Algorithm Alerts Therapists to Suicide Risk in Patients Utilizing the Talkspace Platform (2023) https://www.businesswire.com/news/home/20230912608879/en/Proprietary-AI-Algorithm-Alerts-Therapists-to-Suicide-Risk-in-Patients-Utilizing-the-Talkspace-Platform
- Trade press K-12 Dive, 26M dollar Talkspace contract with NYC stirs student data privacy concerns (2024) https://www.k12dive.com/news/talkspace-nyc-data-privacy-teenspace/727070/
- Investigative Proof News, Woman Talkspace Therapy App Sessions Exposed in Court (2026) https://www.proofnews.org/womans-talkspace-therapy-app-sessions-exposed-in-court/
- Vendor PR Newswire and Universal Health Services, Inc., Universal Health Services, Inc. to Acquire Talkspace, Inc. (2026) https://www.prnewswire.com/news-releases/universal-health-services-inc-to-acquire-talkspace-inc-302708096.html
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Vet connections — Connection authorization
- Store less data — Data minimization
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Escalate checks — State-feedback vigilance
Documented case histories
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Kaiser Permanente Suicide-Risk Model
- Crisis Text Line & Loris.ai
- LyssnCrisis counselor QA at ProtoCall Services (988)
- NarxCare
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- Tessa chatbot replacing the NEDA eating-disorder helpline
- Woebot (a governed app wind-down)