Skip to content

PAN Lab example

SSA 800-Number Conversational AI Assistant

The access gate: a national benefits phone line

A national benefits phone line puts an automated layer in front of every caller. It matches what you say against a fixed set of 74 answers, and it decides whether you reach a person, when, and through which door. Modeled on the United States Social Security Administration's 800-number deployment. In one year automation went from about 300,000 handled calls a month to about 2.9 million, the headline wait fell to seven minutes, and the agency's own inspector general checked the arithmetic and found it correct. The same audit set out what the arithmetic covers: a caller who accepts a callback is counted as a zero wait, and about 25 million calls a year that ended in a hang-up or a busy signal are outside the count entirely. Callers who declined callbacks held for a monthly average that ran from 19 minutes to 100. What this board turns on is that the platform running the access gate also computes the published figures that grade it, so the counting conventions are the finding, and the failure it can produce is not a wrong decision about you but no contact with anyone at all. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. The Lab offers this deployment every tool its own record supports, and applying all of them at once costs about three and a half times the budget and still leaves pathways open, so no amount of money reaches the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the SSA-800-number-class national telephone access gate network: 11 components and 21 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 11 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the national telephone access-gate pattern documented in the Social Security Administration (SSA) 800-number case file, not a reconstruction of the actual system. The gate is drawn as a constrained answer matcher over 74 fixed frequently asked questions, following the inspector general's own description and the word-matching account in the reporting, and not as a generative system; one later article describes it as generative and that characterization is not adopted here.

  • baseline

    Demand reads 3 from documented load, not from a default: 68 million callers served in fiscal 2025, a 65 percent increase over the prior year, between 4.2 and 7.7 million callers a month, about 25 million calls a year ending with no service at all, and a busy rate that reached 29.4 percent in March 2025 when the Fairness Act brought 3.2 million affected beneficiaries onto the line.

  • baseline

    Capacity reads 1 because the human channel is documented as strained and shrinking on both sides at once. The workforce fell about 10 percent net, from roughly 57,000 to roughly 51,400, with about 6,200 departures reported to lawmakers, and about 1,000 field office employees were reassigned onto 800-number duty to hold the line. The announced plan was a larger cut of roughly 12 percent, which is a plan figure and not the net change. On the predecessor platform the fiscal 2024 peak average answer time was 42.4 minutes and an audit tied that platform's unmet requirements to increased waits and to disconnected or unanswered calls. This automation was layered onto a channel that was already failing to answer the phone, which is the opposite of the shape where automation displaces a working manual process.

  • baseline

    The two record stores exist because the gap between them is the case. The platform's operational record holds what happened on each call; the published metric set is computed from it under conventions the audit set out: a caller who accepts a callback counts as a zero wait, and about 25 million calls a year that ended in an abandonment or a busy signal are outside the count. The audit found the published figures accurate as computed. Both halves of that finding travel together everywhere in this network, and neither is stated alone.

  • baseline

    The reconciliation between the two stores is drawn at the low rung rather than at zero, and the distinction is evidence-forced. A variance process genuinely exists and genuinely fired: agency subject-matter experts compare metric reports against raw platform data and meet the platform vendor to resolve differences, and the interactive-voice-response counting method was corrected in April 2025, after which the audit found the metrics accurately reported. What that process checks is arithmetic. The reconciliation that is drawn at zero is the other one, between the disposition layer and the record: an unreturned callback is canceled at end of day and not requeued, and no published figure carries the roughly 25 million calls that ended with no service back into the account of how the line performed.

  • baseline

    The enforcement node is drawn because a downstream action system is documented, not because the shape wanted one. Queue depth passes automatically against thresholds the agency sets into busy messages and ended calls: the prior platform allowed about 12,000 calls across all queues before busy conditions, and the current platform lets the agency set the number, in some queues up to 40,000. That makes denial of access a parameter someone chooses rather than a consequence nobody controls, which is why it is drawn as structure and not as prose.

  • baseline

    The spoken handoff is drawn as a bounded automated screen at the low rung. It is documented to exist, since a caller can say agent to ask for a person, and it is documented to work only some of the time, since complaint records describe requests that produced no handoff and calls the system ended instead. No published figure states how often the route works, so the rung reflects a screen that is partial by construction rather than a measured catch rate.

  • baseline

    The two reviewer nodes are wired in by different pathways, and that asymmetry is the finding rather than a drafting choice. The audit reached the record: its team obtained raw platform data and checked each published figure against it. The congressional channel reached the gate, but only through constituent complaint and press account, since no published audit has put a rate on the gate's misrouting or on calls it ended. That is why the model-side independent check is drawn at zero while the record-side variance review is drawn present.

  • baseline

    The strongest operator-to-model pathway in this network is deployment timing, and it is documented rather than inferred. The gate was built and tested under one administration and held back because the team wanted the automation to give consistent and accurate answers first; it was deployed by the next administration in April 2025 and pushed toward about 1,200 field offices by August 2025. A separate automated phone check was added in April 2025 and withdrawn in May. The predecessor platform was launched and abandoned inside about ten months. Deployment state moved with leadership across five commissioners or acting commissioners in one fiscal year, and no published evaluation result marks any of those turns. The late-May 2025 statement to employees that AI would take wait times down to single digits is carried on this same pathway, because its documented consequence is the deployment push and the record documents no change in floor practice that followed from the statement itself; a separate pathway from leadership to the floor is not drawn.

  • baseline

    The separate phone anti-fraud check of April to May 2025 is a different tool and is not drawn. It flagged 2 of more than 110,000 claims and slowed retirement claim processing by 25 percent before its three-day holds were removed. It is cited here only as evidence that this agency holds and exercises the authority to stop an automated channel on this line, and its figures are never merged with the answer gate's.

  • baseline

    No egress node is drawn, and the omission is deliberate. The platform is vendor-operated, the agency's own experts must meet the vendor to resolve metric variances, and the vendor of the current platform is not named in the public audit and is not asserted anywhere here. What the record does not document is any data-protection finding, breach, or ungoverned crossing, so an external-boundary node would be decorative. The 160 million dollar figure that appears in this record belongs to the abandoned predecessor platform and never to the current one.

  • assumed

    The frontline is drawn as one operator class. The record documents that about 1,000 field office employees were moved onto 800-number duty from July 2025 and that all 1,231 field offices were brought behind the same gate on the same platform, and it documents no differential adoption or error between the teleservice representatives and the reassigned staff: the same handed-on calls reach both, and both write to the same record. Their documented difference is share, about 1,000 people against the national teleservice workforce, working a fraction of the record and of the callback list alongside their own offices' business; that is an intensity, carried in the class's own copy rather than as a second wiring. A pathway for practice spreading between the two groups is not drawn, because the record documents the unification and the reassignment and documents no rate of spread; the earlier drawing disclosed that rate as a modeling convention, and this drawing carries the same documented facts without it.

  • baseline

    One fixed set of 74 answers is matched for every caller in the country, and in fiscal 2025 the same gate was extended behind all 1,231 field offices, so a gap in one answer repeats nationally rather than case by case. The model self-loop carries that correlated reach at the top rung, which is higher than this catalogue's usual monoculture reading and is justified by a single national answer set rather than a shared vendor product.

  • assumed

    The call record and transcript pathway is marked as carrying sensitive payload because the flow itself does: identifiable claim and benefit conversation, written automatically at the rate of about 68 million calls a year with no person confirming any entry, plus retained interaction transcripts the agency reviews. The record documents no data-protection finding, breach, or complaint about this retention, so the mark reflects what the pathway carries and not an established governance failure.

  • assumed

    Served people are not in the dynamics. The 74 million beneficiaries this line exists for, the caller who was told about railroad retirement when asking about disability, the caller who waited an average of 100.1 minutes in January 2025 after declining a callback, and every one of the roughly 25 million calls a year that ended with no service are boundary quantities recorded in the case file. This diagram propagates institutional error through the operator network, estimates no differential harm, and computes no caller outcome from anything drawn here.

What this example does not show

  • The metrics dispute here is DEFINITIONAL, and both readings come from the same document. The inspector general found the published telephone figures accurate as computed. The same audit documented that the headline average counts a caller who accepts a callback as a zero wait, and that about 25 million calls a year ending in a hang-up or a busy signal are outside the count. Any use of the seven-minute figure carries that caveat, and neither the vindication reading nor the misleading reading is stated on its own.
  • The gate is modeled as a constrained matcher over 74 fixed answers, not as a generative system. The inspector general describes a conversational question-and-answer chatbot over 74 frequently asked questions and one outlet describes word matching; a later article calls it generative AI that learns as it goes, which tracks agency and vendor framing and is not adopted here. No platform vendor is asserted for the current system: the public audit refers only to the telecommunications platform vendor. The figure of more than 160 million dollars belongs to the abandoned predecessor programme, not to the system on this diagram.
  • The gate's own failure rate is not established. Misrouting, wrong-topic answers, refusals to hand off and calls ended early are documented through caller complaints, senator letters and journalism; no audit has published a misrouting or wrongful-disconnection rate. The quantified failure edges on this diagram, the roughly 25 million unserved calls and the busy-rate spikes, are platform-level and are not attributed to the chatbot. The separate phone anti-fraud check of April to May 2025 is a different tool and its figures never merge with these.
  • The workforce reduction is about 10 percent net to date, from roughly 57,000 to roughly 51,400; the roughly 12 percent figure is the announced plan and is not the change that has happened. The shelve-then-revive arc and the withdrawal of the separate anti-fraud check are authority actions by successive administrations, not adjudicated findings that any system failed. The latest primary document here is dated December 22, 2025; agency statements from June 2026 are claims.
  • Served people are not modeled here. The 74 million beneficiaries this line exists for, the caller told about railroad retirement when asking about disability, and every one of the roughly 25 million calls a year that ended with no service are boundary quantities recorded in the case file. The Lab models institutional propagation through the operator network, estimates no differential harm to served people, and computes no caller outcome from anything on this diagram.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The Social Security Administration deployed a conversational question-and-answer chatbot on its national 800-number in April 2025, answering 74 frequently asked questions before a caller reaches an employee. Automation on that line went from roughly 300,000 handled calls a month in fiscal 2024 to roughly 2.9 million a month in fiscal 2025, peaking at 5.1 million automated calls in March 2025, while the agency served 68 million callers, a 65 percent increase over the prior year, with a workforce that fell about 10 percent net from roughly 57,000 to roughly 51,400 and about 1,000 field office employees reassigned onto 800-number duty. About 25 million fiscal 2025 calls ended with no service, abandoned in queue or met with a busy signal, and the busy rate spiked to 29.4 percent in March 2025 during the Social Security Fairness Act surge, which affected 3.2 million beneficiaries.

    empirical
    • Government SSA Office of the Inspector General, The Social Security Administration's Telephone Metrics (Audit Report 032517) (2025) https://oig.ssa.gov/assets/uploads/032517.pdf
    • Investigative KFF Health News (Darius Tahir), Social Security Praises Its New Chatbot. Ex-Officials Say It Was Tested but Shelved Under Biden (2025) https://kffhealthnews.org/aging/social-security-chatbot-customer-complaints-glitches/
    • Investigative AARP, How AI is Changing Social Security Customer Service (2025) https://www.aarp.org/social-security/ai-customer-service/
  • The agency's inspector general found the published telephone metrics arithmetically accurate as computed and documented what they cover: the headline average speed of answer, 12.7 minutes in October 2024, a 29.7 minute peak in January 2025 and 7.0 minutes in September 2025, counts a caller who accepts a callback as a zero wait, and the roughly 25 million calls a year ending in a hang-up or a busy signal are excluded. The 23.8 million callers who accepted callbacks in fiscal 2025 waited an average of 61.5 to 151.8 minutes depending on the month, and the 9.3 million who declined and held waited an average of 18.8 to 100.1 minutes. A separate 87 percent first-contact-resolution figure comes from a post-call survey that is not offered to callers served only by automation, and no audit has published a misrouting or wrongful-disconnection rate for the chatbot itself.

    empirical
    • Government SSA Office of the Inspector General, The Social Security Administration's Telephone Metrics (Audit Report 032517) (2025) https://oig.ssa.gov/assets/uploads/032517.pdf
    • Investigative Nextgov/FCW, SSA phone wait times longer than publicly reported metrics, per OIG report (2025) https://www.nextgov.com/digital-government/2025/12/ssa-phone-wait-times-longer-publicly-reported-metrics-oig-report/410360/
  • Deployment decisions on this line have followed changes of leadership rather than published evaluation results. The chatbot was developed and tested under one administration and shelved as not ready, with the former chief information officer saying the team wanted to ensure the automation produced consistent and accurate answers and that this would take more time; the next administration deployed it in April 2025 and targeted extension to roughly 1,200 field offices by August 2025. A separate phone anti-fraud AI check was added in April 2025 and its three-day claim holds removed in mid-May 2025 after it flagged 2 of more than 110,000 claims while slowing retirement claim processing by 25 percent, a different tool whose figures do not merge with the chatbot's. The predecessor Next Generation Telephony Project, built by Verizon Business Network Services under a February 2020 contract, was abandoned on August 22, 2024 after about ten months and more than 160 million dollars paid, with an audit finding that contract lacked performance-based quality standards; the vendor of the current cloud platform is not named in the public audit. Five commissioners or acting commissioners served during fiscal 2025, and the metrics published on the agency's performance website were added and removed according to what each leadership believed were the most important metrics for the public.

    empirical
    • Investigative KFF Health News (Darius Tahir), Social Security Praises Its New Chatbot. Ex-Officials Say It Was Tested but Shelved Under Biden (2025) https://kffhealthnews.org/aging/social-security-chatbot-customer-complaints-glitches/
    • Investigative Nextgov/FCW, SSA changes phone fraud policies after finding very little fraud (2025) https://www.nextgov.com/digital-government/2025/05/ssa-changes-phone-fraud-policies-after-finding-very-little-fraud/405380/
    • Government Office of Senator Elizabeth Warren, Warren, Wyden, Sanders, Gillibrand Demand Answers on Reckless AI Tool Rollout at SSA (letter of June 24, 2025, released July 1, 2025) (2025) https://www.warren.senate.gov/newsroom/press-releases/warren-wyden-sanders-gillibrand-demand-answers-on-reckless-ai-tool-rollout-at-ssa
    • Government SSA Office of the Inspector General, SSA Abandoned 160 Million Dollars Plus Next Generation Telephony Project (news release) (2025) https://oig.ssa.gov/news-releases/2025-04-23-ssa-abandoned-160-million-next-generation-telephony-project/
    • Government SSA Office of the Inspector General, The Social Security Administration's Telephone Metrics (Audit Report 032517) (2025) https://oig.ssa.gov/assets/uploads/032517.pdf
  • A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.

    empirical
    • Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
    • Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories