PAN Lab example
NYC MyCity business chatbot
Exposure is not correction: a public-facing government advice chatbot
This chatbot answered business owners' questions about city rules in a confident, official-sounding voice — and it was repeatably wrong on the law, telling businesses they could take a cut of workers' tips or that landlords need not accept housing vouchers. Modeled on the NYC MyCity Business chatbot. The catch is not that a model erred; every model errs. It is that the person acting on the answer was the public, with no caseworker to check it, so the legal risk fell on third parties who were never in the conversation — and that the error was published, then formally audited, and the tool kept running for roughly two years until a budget cut, not an accuracy fix, ended it. On this shape the leverage is all upstream: there is no operator to train and no verify habit to protect, so you gate the answer against the law, label what it is, and hold a take-down trigger — and you make the audit that exposed it actually bind.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Public-adviser-class generative chatbot with rejected external oversight network: 5 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 4 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the public-facing generative-adviser pattern documented in the NYC MyCity Business chatbot case file — not a reconstruction of the actual chatbot. It is deliberately the reckless-pole counterpart to the library's two verify-before-use benefits copilots (the Nava-class and Benefit-Navigator-class assistive chatbots): those place a professional between the model and the affected person and their whole safety case is a verify-before-use habit, whereas here the operator is the untrained public, so there is no verify habit to protect and the harm externalizes onto third parties.
- baseline
The model-to-operator deference channel is drawn at its maximum because the public acted directly on answers delivered in an authoritative municipal voice, with no caseworker or professional intermediary to verify before acting. This is the structural inverse of the sibling copilots, where that channel is bounded by a caseworker who reads the citation before relaying — the single design difference that separates the cautious pole from the reckless one on the same underlying technology.
- baseline
This shape's defining feature is an absence: the model-side inhibiting check is drawn empty at baseline, because no independent legal-truth check gated a generated answer against statute or an authoritative rules engine before it reached the public, which is how confident answers that would break the law (on tips, on Section 8 and source-of-income housing rules, on cashless-business and lockout rules) reached business owners as apparent official guidance.
- baseline
The model-to-model self-loop encodes correlated error by construction: one public-facing model answered everyone identically, so a confidently wrong answer was not idiosyncratic but was served to every person who asked. The Markup's ten staffers all received the same wrong answer that landlords need not accept vouchers; the documented inconsistency was that the same question returned different answers over time and was sensitive to trivial input variation, not that the ten testers disagreed with each other.
- assumed
The disclaimer-and-scope-redirect node is drawn as a porous named safeguard carrying no effective correction flow, reflecting the documented account that the only user-facing safeguard was a beta-product disclaimer telling users to double-check via links, that the later patch narrowed scope rather than fixing accuracy, and that the bot itself, when asked, still said it could be used for professional business advice. The weak 'double-check the links' pathway is carried by the store-to-operator edge, drawn low.
- baseline
The external-oversight node is drawn as an inflow whose correction pathway runs latent at baseline. The investigative press, an incident database, and the December 2025 Comptroller performance audit surfaced the errors, but the audit's seven recommendations, including AI red-teaming, were rejected, and the chatbot ran for roughly two years until it was shut down as a budget cut rather than an accuracy fix. The case turns on this: an oversight signal that lands but is rejected is not a control.
- assumed
The model's raw error rate is a modeling choice, not a measured per-interaction rate. The documented illegal-advice character (confidently and repeatably wrong on legal obligations) is qualitative, based on specific tested questions; the only quantitative accuracy-style figures come from the 2025 audit's own sample and testing (a household-of-70 feedback share and a failure-to-answer count from the owner's internal weekly report), not a lifetime error rate, and the city's counter-claims of 'thousands of accurate answers' and hallucinations 'down exponentially' were official and unquantified.
- assumed
Served and affected people are not in the dynamics. The workers denied tips and the tenants refused vouchers or locked out are third parties who were never in the conversation and who bore the downstream legal risk of a wrong answer; this Lab models institutional propagation only, and that externalized harm is documented in the case file and recorded outside any diagram like this one. The separate MyCity childcare eligibility-automation component and its figures are a distinct system and are not modeled or attributed here.
What this example does not show
- The people actually harmed — the workers denied tips, the tenants refused vouchers or locked out — are third parties who were never in the conversation and are not modeled here; the Lab models institutional propagation only, and those downstream legal harms are documented in the case file and measured outside any diagram like this one.
- The 2024 findings that the chatbot was confidently and repeatably wrong are qualitative, based on specific tested questions rather than a sampled error rate. The only quantitative accuracy-style figures come from the December 2025 Comptroller audit's own sample and testing — a 71.4% negative share among the 70 users who left feedback (50 of 70), which the city disputes as roughly 2.25% of all responses, and a count of 23 of 48 tested government questions unanswered drawn from the owner's internal weekly report — not a lifetime per-interaction error rate.
- This scenario models the public-facing advice chatbot only. The separate MyCity childcare eligibility-automation component, and figures such as its application-ineligibility rate, belong to a distinct system and are neither modeled nor attributed to the chatbot here; the wider MyCity program's 100-million-dollar cost is the whole system, not the chatbot alone.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
New York City launched the MyCity Business chatbot in 2023 on Microsoft Azure AI as a public-facing generative-AI adviser for business owners. A March 29, 2024 investigation by The Markup with THE CITY and Documented NY found it confidently and repeatably wrong on legal obligations, advising businesses in ways that would break the law, including that employers could take a cut of workers' tips, that landlords need not accept Section 8 vouchers or source-of-income tenants (illegal in New York City), that stores could go cashless against a 2020 city law, and that funeral-price disclosure could be concealed against the federal funeral rule; when ten staffers asked the housing-voucher question they received the same wrong answer, which had changed from an earlier correct one, showing the tool was non-deterministic. The 2024 findings are qualitative, based on specific tested questions rather than a sampled error rate. The city relabeled the tool a beta product with a disclaimer and applied a scope-narrowing patch rather than withdrawing it, kept it online for roughly two years, and shut it down in early 2026 as a budget cut rather than an accuracy fix.
empirical- Investigative The Markup, NYC's AI chatbot tells businesses to break the law (2024); OECD AI incident https://themarkup.org/artificial-intelligence/2024/03/29/nycs-ai-chatbot-tells-businesses-to-break-the-law
- Investigative The Markup and THE CITY, Malfunctioning NYC AI Chatbot Still Active Despite Widespread Evidence It's Encouraging Illegal Behavior (2024) https://themarkup.org/artificial-intelligence/2024/04/02/malfunctioning-nyc-ai-chatbot-still-active-despite-widespread-evidence-its-encouraging-illegal-behavior
- Investigative Reuters (Jonathan Allen), New York City defends AI chatbot that advised entrepreneurs to break laws (2024) https://finance.yahoo.com/news/1-york-city-defends-ai-011454323.html
- Investigative The Markup (Colin Lecher and Katie Honan), Mamdani to Kill the NYC AI Chatbot We Caught Telling Businesses to Break the Law (2026) https://themarkup.org/artificial-intelligence/2026/01/30/mamdani-to-kill-the-nyc-ai-chatbot-we-caught-telling-businesses-to-break-the-law
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
A December 30, 2025 performance audit of the MyCity system, issued under New York City Comptroller Brad Lander, found the chatbot 'appears to be unable to provide accurate or consistent information' and reported that the wider MyCity system had cost over 100 million dollars across more than 120 agreements with about 50 vendors, lacked a system development plan, and had not delivered the promised single-form access to city benefits; the Office of Technology and Innovation disagreed with all seven of the audit's recommendations, including one to conduct AI red-teaming. Among the audit's figures, an internal weekly production report reproduced in the audit showed the chatbot did not answer 23 of 48 tested government questions, and of the more than 2,200 questions asked in July and August 2025 the 70 users who left thumbs-up-or-down feedback were 71.4 percent negative (50 of 70), a share the city disputes as roughly 2.25 percent of all responses, with the audit rebutting that denominator. The 100-million-dollar figure is the whole MyCity system, not the chatbot alone.
empirical- Government evaluation Office of the New York City Comptroller (Brad Lander), Audit Report on the New York City Office of Technology and Innovation's MyCity System (2025) https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/
Where this connects
Institutional pressures in this domain
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Public benefits & eligibility domain page.
Levers available here and the patterns behind them
- Check with a second model — Cross-model verification
- Mark AI-written records — Provenance labeling
- Pause AI on alarms — Deployment circuit-breaker
- Require sign-off — Conformity assessment gate
- Review on schedule — Oversight cadence & retrospectives
- Gate vendor updates — Vendor quality gate
- Keep prompts neutral — Framing and mirroring reduction
- Upgrade model — Improve the model
- Escalate checks — State-feedback vigilance
Documented case histories
- NYC MyCity business chatbot
- Michigan MiDAS
- Robodebt (Australia)
- Indiana / IBM eligibility modernization
- Rotterdam welfare-fraud risk model
- Arkansas ARChoices / ARIA
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- SyRI (Netherlands)
- CNAF benefit-fraud risk score (France)
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Udbetaling Danmark data-driven control (Denmark)
- BOSCO (Spain)
- Serbia Social Card (Socijalna karta)
- UK DWP Universal Credit Advances fraud model
- ID.me identity verification as an unemployment eligibility gate
- Medicaid unwinding: automated ex parte renewal at population scale
- INSS auto-analysis: when the productivity metric makes denial the fastest way out
- Samagra Vedika
- Workforce Australia Targeted Compliance Framework: automated payment sanctioning after Robodebt
- Nevada DETR generative-AI unemployment appeals
- Tennessee TennCare TEDS