Domain Atlas / Benefits navigation & public-facing chat
SSA 800-Number Conversational AI Assistant
The Social Security Administration deployed a conversational question-and-answer chatbot on its national 800-number in April 2025, answering 74 frequently asked questions before a caller reaches an employee. Automation on that line went from roughly 300,000 handled calls a month in fiscal 2024 to roughly 2.9 million a month in fiscal 2025, peaking at 5.1 million automated calls in March 2025, while the agency served 68 million callers, a 65 percent increase over the prior year, with a workforce that fell about 10 percent net from roughly 57,000 to roughly 51,400 and about 1,000 field office employees reassigned onto 800-number duty. About 25 million fiscal 2025 calls ended with no service, abandoned in queue or met with a busy signal, and the busy rate spiked to 29.4 percent in March 2025 during the Social Security Fairness Act surge, which affected 3.2 million beneficiaries.[3]
What happened
The Social Security Administration's national 800-number is one of the largest public contact channels in the United States, and in fiscal year 2025 it served 68 million callers, a 65 percent increase over the year before, with monthly callers ranging from 4.2 million in November 2024 to 7.7 million in March 2025. In April 2025 the agency rolled out a conversational question-and-answer chatbot on that line: an automated layer that matches what a caller says against 74 frequently asked questions and answers them before an employee is reached, working alongside round-the-clock interactive self-service for tasks such as benefit verification, Medicare card replacement, claim status and form requests. Press descriptions of the technology conflict. One outlet in July 2025 described a word-matching, non-generative question-and-answer bot; another in October 2025 described generative AI that learns as it goes, which tracks agency and vendor framing. The inspector general describes only a conversational question-and-answer chatbot over 74 frequently asked questions, and that constrained description is the one used here.
The volumes moved fast. Automation handled roughly 300,000 calls a month in fiscal 2024 and roughly 2.9 million a month in fiscal 2025, peaking at 5.1 million automated calls in March 2025; about 41 percent of calls were reported handled by the chatbot as of July 2025, and 1.6 million of about 5.1 million September 2025 calls resolved through automated self-service. The context was a shrinking workforce: agency staffing fell from about 57,000 to about 51,400, a net reduction of about 10 percent, with about 6,200 departures the commissioner reported to lawmakers in June 2025, and about 1,000 field office employees reassigned onto 800-number duty; the announced plan was a larger reduction of roughly 12 percent, which is a plan figure rather than the change to date. At an employee all-hands in late May 2025 the commissioner said the agency would get wait times down to single digits using AI. Separately, before lawmakers in June 2025, he described an all-time staffing low and an all-time technological high. Those were two statements to two audiences and they are kept apart here.
The published performance figures improved. The headline average speed of answer was 12.7 minutes in October 2024, peaked at 29.7 minutes in January 2025, and reached 7.0 minutes in September 2025, against a fiscal 2024 peak of 42.4 minutes in November 2023. Senator Elizabeth Warren wrote to the acting inspector general on July 24, 2025 alleging the published metrics were misleading and incomplete and citing a staff survey that found average waits of about an hour and 45 minutes; the commissioner had agreed to an investigation in a meeting the day before, the watchdog opened the review in September 2025, and audit report 032517 was published on December 22, 2025. The audit found the published metrics arithmetically accurate as computed. It also set out what they cover. A caller who accepts a callback is counted as a zero wait, which lowers the average: 23.8 million callers accepted a callback in fiscal 2025 and waited an average of 109.4 minutes in October 2024, 151.8 minutes at the January 2025 peak and 61.5 minutes in September 2025, none of which enters the headline. The 9.3 million callers who declined callbacks and held waited an average of 51.2 minutes in October 2024, 100.1 minutes in January 2025 and 18.8 minutes in September 2025. About 25 million fiscal 2025 calls ended with no service at all, abandoned in queue or met with a busy signal, and are excluded from the headline; callers waited an average of 22 to 38 minutes before hanging up, and unreturned callbacks are canceled at the end of the day and are not returned to the queue. The busy or polite-disconnect rate averaged about 6 percent across fiscal 2025 and spiked to 29.4 percent in March 2025 and 15.4 percent in April 2025 during the Social Security Fairness Act call surge, an Act that affected 3.2 million beneficiaries. Those thresholds are set: the prior platform allowed roughly 12,000 calls across all queues before busy conditions, while the current platform lets the agency choose, in some queues up to 40,000. A separate first-contact-resolution figure, 87 percent of post-call survey respondents in fiscal 2025, comes from a survey that is not offered to callers served only by automation.
Documented failures of the chatbot itself are qualitative. Reported modes include answering a different question than the one asked, inaccurate answers on spousal and survivor benefits, confusing retirement benefits with Supplemental Security Income, refusing to hand a caller off to an employee, and ending calls it treated as resolved; one disability claimant instead received information about railroad retirement and domestic partnerships. Callers can say agent to ask for a person, but the same complaint record shows requests that produced no handoff. No audit has published a misrouting or wrongful-disconnection rate for the bot, so the quantified failure edges in this record are platform-level rather than bot-attributed. On June 24, 2025, in a letter released July 1, Senators Warren, Wyden, Sanders and Gillibrand demanded answers on what they called a reckless AI rollout, noting the chatbot was deployed with little consultation with Congress, advocates, or other key stakeholders, and that it came a month after the agency reversed a separate phone anti-fraud AI check which had flagged 2 of more than 110,000 claims as high-probability fraud while slowing retirement claim processing by 25 percent; its three-day claim holds were removed in mid-May 2025. That anti-fraud tool is a different system and its figures are never merged with the chatbot's. A year earlier the posture had been bipartisan and procedural: on August 6, 2024, Finance Committee chairman Wyden and ranking member Crapo jointly asked how the agency's dozen-plus AI systems complied with OMB M-24-10 and the National Institute of Standards and Technology (NIST) AI Risk Management Framework, and how human discretion was preserved.
The chatbot was not new when it shipped. It was developed and tested under the previous administration and shelved as not ready; the former chief information officer said the team wanted to ensure the automation produced consistent and accurate answers, which was going to take more time. The next administration deployed it in April 2025 and described it as a success, and by August 2025 it was targeted for extension to roughly 1,200 field offices. This channel has a history of reversals: the agency paid Verizon Business Network Services more than 160 million dollars for the Next Generation Telephony Project under a February 2020 contract, ran the 800-number on it from November 2023, and abandoned it on August 22, 2024 after about ten months, moving to a new cloud platform whose vendor the public audit does not name. An April 2025 audit found the abandoned contract lacked performance-based quality standards and tied its unmet requirements to increased wait times and to disconnected or unanswered calls. During fiscal 2025 the agency had five commissioners or acting commissioners, and the telephone metrics published on its performance website were added and removed according to what each leadership believed were the most important metrics for the public; the agency corrected its interactive-voice-response counting methodology in April 2025, so the metrics were accurately reported from May 2025 onward. The agency's December 22, 2025 press release said the inspector general report confirmed significant customer service improvements, quoting the commissioner that the agency was serving more Americans at significantly faster speeds than ever before, and in June 10, 2026 congressional testimony he described record-low helpline wait times, with a representative raising the callback convention and the commissioner calling it an industry standard.
The sociotechnical reading
Most AI systems in this atlas decide something about a person: a score, a rank, an eligibility. This one decides whether the person reaches anyone at all. That single difference reorganizes everything downstream. The failure surface is not a wrong judgment about a claim but misrouting, a wrong-topic answer, a call ended early, and an engineered busy signal, and it falls across the whole population a national line serves rather than the subset a risk model flags. It also means the harm is close to invisible from inside: a caller who gave up at minute 26, or whose callback was cleared at end of day, generates no case, no appeal, and no artifact anyone has to read.
The second mechanism is what makes this case unusual as governance evidence. The same platform that runs the gate computes the figures that grade the gate. That is not an allegation of manipulation, and the audit is precise on the point: the published numbers are accurate as computed. The problem is definitional and it is structural. A convention that counts a caller who accepts a callback as a zero wait removes 23.8 million people's waiting time, averaging over an hour and a half at the peak, from the headline. A convention that excludes calls ending in a hang-up or a busy signal removes about 25 million contacts a year, which is to say it removes precisely the population the gate failed. A satisfaction survey offered only to callers who reached a person cannot register anything the automation did wrong. And a callback nobody returned is canceled overnight rather than requeued, so the evidence of that failure is not merely uncounted, it is deleted daily. None of those choices requires bad faith. Together they produce a measurement system that is structurally blind in the exact direction the deployment is most likely to fail, and the improving number then travels upward, gets cited publicly, and justifies deeper automation and deeper staffing cuts. Both readings of the audit, that the agency was vindicated and that the agency was misleading, come from the same document, which is why the arithmetic finding and the definitional caveats have to travel together every time either is used.
What the map draws is where the checking is and is not. Oversight here is real and it has teeth: two completed audits, a wait-time review a senator demanded and a commissioner agreed to, bipartisan questioning on AI governance framework compliance, and at least three occasions where scrutiny changed behavior, including a separate automated phone check that was withdrawn within a month of drawing attention. It is also aimed at one thing. The audit reached the record; it did not reach the answers. There is a reconciliation behind the published figures, and it works, and it corrected a counting method in April 2025. There is no reconciliation at all behind the disposition layer that decides who gets a busy signal, and there is no independent read of what the 74 answers actually tell callers. The gap between those two facts is this network's shape. A third pattern runs underneath both: deployment state moves with political leadership rather than with evaluation results. A system judged not ready by the people who built it was shipped by the next administration; a fraud check was added in April and withdrawn in May; a 160 million dollar platform was launched and abandoned inside ten months; five commissioners served in one fiscal year and the published metric set changed with them. The lesson for the Field Guide is that when access itself is the thing being automated, the decisive governance artifact is not an accuracy threshold but a counting rule, because whoever defines what a served call is has already decided what the system is allowed to fail at. Nothing in this record measures what callers experienced, and the effect on the 74 million beneficiaries at the other end of this line was never a quantity any part of this system computed.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library.