Skip to content

PAN Lab example

IRS collection chatbots

Expanded without a ruler: a federal collection chatbot with no performance measures

A federal collection agency built a chatbot and a live-chat line to steer people off its phone queue, then expanded them and made them permanent - without ever measuring whether they worked. Modeled on the IRS Automated Collection System chat applications - their shape, not the real tool. The chatbot decides nothing; it is scripted, not generative. The danger is not a wrong answer, it is an unmeasured one: the statistics the agency did collect were so unreliable they showed one worker handling 603 chats at once against a system cap of three, while assistors juggling concurrent authenticated chats risked showing one taxpayer's account to another. The only thing that caught the broken numbers was an outside audit. Watch the feedback link that is drawn on the map but severed in practice - and the two checks that were never switched on.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the IRS ACS chatbot-class collection-deflection funnel network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the monitoring-absence pattern documented in the IRS ACS chat-applications case file - not a reconstruction of the actual system. The atlas-relevant object is not a decision the system makes but the measurement it does not make: the ACS chatbot is scripted and explicitly non-generative and makes no eligibility or scoring decision, and the defining finding is that the IRS had no performance measures for the program at all while expanding it and making it permanent. Any reading that implies algorithmic decision-making, or a generative model, misreads it.

  • baseline

    The corrupt monitoring store is the payload: the platform's statistical reports are unreliable (a miscalculated four-hour handle-time metric inflated concurrency to as many as 603 chats at once against a systemic cap of 3, over 16,000 records showed potential concurrency, and 22,628 resolution codes exceeded the chats they should match), and management read them, knew the counts did not reconcile, and did not investigate. Its high write-contamination and near-zero decontamination on the map encode that a feedback link formally present but functionally severed is worse than a visible blank, because its readings launder an ungoverned expansion as measured success.

  • baseline

    The only functioning oversight is an external, episodic inspector-general audit, drawn as a strong check that sits entirely outside the org, while the internal evaluative-review check is drawn empty. That contrast is the case's distinctive safety shape: unlike the domain's staged-pilot-gate government chatbot, whose evaluation team held a version below its accuracy bar before release, here the internal monitoring channel is severed and the audit is the only thing that bites - once in years, and forced to discard most of the data it was auditing.

  • baseline

    A second dynamic couples operator load to disclosure risk: of a judgmental sample of 40 live assistors, 24 worked multiple chats concurrently and 12 of those had at least one authenticated chat open while working another, raising the documented risk of disclosing one taxpayer's information into another taxpayer's session - carried on the privacy-sensitive assistor-to-records write and records-to-assistor read. A 20% assistor cut (159 to 127) inside an agency-wide 25% workforce cut raises per-operator load while the measurement channel stays broken, so the org cannot even see the overload rising. These are institutional signals about a record pathway; no disclosure to any actual taxpayer is asserted or computed.

  • assumed

    The 603-concurrent and 54%-unresolved figures are computed from data TIGTA itself deems unreliable; TIGTA states the precise number of unresolved chats cannot be determined, and the 60% concurrency rate comes from a judgmental (nonprobability) sample of 40 assistors that cannot be projected to the 253-assistor population - TIGTA could hand-verify only a 1.20-to-3.88 concurrency band. Voice bots were excluded from audit testing, so the error rates apply to the chat applications only, and the 2022 volume and dollar figures (450,000-plus interactions, 4.8 million calls, 7,600 installment agreements, over 50 million dollars) are unaudited agency self-claims from a promotional essay now marked historical. The vendor is unnamed in every fetched source; naming one would be invention.

  • assumed

    Served people, and the tax obligations they do or do not resolve, are not in the dynamics; this Lab reads institutional propagation only. The Spanish-language chatbot lacks the free-text search box entirely, limiting Spanish speakers to prepopulated options - a differential-access concern recorded honestly here and never computed as a demographic outcome. The 54% of chats that go unresolved (with abandonment tripling year over year) recirculate demand back toward the toll-free phone line the funnel was built to relieve, which is why the containment claim justifying expansion could not be substantiated. A concurrency record, a resolution code, or a disclosure-risk signal on this map is an institutional signal, never a person, and a safe starting baseline is a property of this model, not a safety promise for any real deployment.

What this example does not show

  • Served people - the taxpayers who use the chatbot, the live chat, and the voice bots - are not modeled here; the Lab reads institutional propagation only. The concurrency-driven disclosure risk (an assistor surfacing one taxpayer's account in another's authenticated session) is an institutional signal about a record pathway, never a computed disclosure to any actual person, and the Spanish-language chatbot's missing free-text search box is a differential-access concern recorded honestly, not a demographic outcome computed from anything here.
  • The 603-concurrent and 54%-unresolved figures are computed from data the inspector-general audit itself deems unreliable; the audit states the precise number of unresolved chats cannot be determined, and the 60% concurrency rate comes from a judgmental (nonprobability) sample of 40 assistors that cannot be projected to the 253-assistor population - only a 1.20-to-3.88 concurrency band was hand-verified. These are treated as directional evidence of a monitoring failure, not calibrated quantities.
  • Voice bots were excluded from the audit's testing, so the chatbot content-deficiency findings (14% of process flows, 83% of tested keywords) apply to the chat applications only and are not extended to the voice channel. The 2022 volume and dollar figures (450,000-plus chatbot interactions, 4.8 million and 1 million-plus voice-bot calls, 7,600 installment agreements covering over 50 million dollars) are unaudited agency self-claims from a promotional essay now marked historical; the audit specifically found the agency could not substantiate its phone-demand-reduction claim.
  • This is a scripted, explicitly non-generative chatbot, not a generative-AI or scoring or eligibility-decision system, and the platform vendor is unnamed in every source. Its atlas relevance is the monitoring-absence topology - a feedback link formally present but functionally severed, with an external episodic audit the only functioning oversight - and the operator-overload disclosure coupling. A safe starting baseline is a property of this model, not a safety promise for any real deployment.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported that the IRS expanded its Automated Collection System chatbot and live-chat program and made live chat permanent while having no performance measures for it, despite a Taxpayer First Act requirement for metrics and benchmarks, and that management's claim the bots reduced telephone demand could not be substantiated; the statistical reports the IRS did collect were deemed unreliable, in one instance showing a single assistor apparently working 603 chats at once against a systemic cap of three, attributed partly to a miscalculated handle-time metric the vendor had not resolved as of December 2025.

    empirical
    • Government evaluation Treasury Inspector General for Tax Administration, Opportunities Exist to Improve the Quality of Chat Applications (Final Audit Report, Report Number 2026-308-029, 2026) https://www.oversight.gov/sites/default/files/documents/reports/2026-06/2026308029fr.pdf
    • Trade press Bracken, IRS live chat apps have room for improvement, watchdog finds (FedScoop, 2026) https://fedscoop.com/irs-live-chat-apps-chatbots-report/
    • Trade press Bramwell, Are IRS Chatbots Really Helping Taxpayers? (CPA Practice Advisor, 2026) https://www.cpapracticeadvisor.com/2026/07/08/are-irs-chatbots-really-helping-taxpayers/186261/
  • In the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple chats concurrently and 12 of those 24 had at least one authenticated chat open while working another, which TIGTA reported as raising the risk of disclosing taxpayer information to the wrong taxpayer; the audit also reported 635,684 resolution codes against 613,056 chats (a mismatch management knew of but did not investigate) and, in March 2025 hand-testing, 14% of chatbot process flows deficient and 83% of tested keywords unrecognized or insufficient, with the figures drawn from a nonprobability sample and data the audit itself characterized as unreliable and not projectable to the full assistor population.

    empirical
    • Government evaluation Treasury Inspector General for Tax Administration, Opportunities Exist to Improve the Quality of Chat Applications (Final Audit Report, Report Number 2026-308-029, 2026) https://www.oversight.gov/sites/default/files/documents/reports/2026-06/2026308029fr.pdf
    • Trade press Bramwell, Are IRS Chatbots Really Helping Taxpayers? (CPA Practice Advisor, 2026) https://www.cpapracticeadvisor.com/2026/07/08/are-irs-chatbots-really-helping-taxpayers/186261/
    • Trade press Cohn, IRS chatbot results may be wrong (Accounting Today, 2026) https://www.accountingtoday.com/news/irs-chatbot-results-may-be-wrong

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories