PAN Lab example
EDD Virtual Assistant
Two levels up: a benefits assistant and its budget
A state benefits agency runs a conversational assistant in two tiers. One is open to anyone, around the clock, in eight languages, and answers general questions about unemployment, disability and paid family leave. The other sits behind a login and tells a signed-in customer their own claim status, payments and eligibility. A person can be reached by chat on weekdays between 9 a.m. and 2 p.m., after identity verification, and the conversation travels with them. Modeled on California's unemployment agency and its modernization programme. What makes this network unusual is where the checking sits. No body oversees the assistant. Oversight attaches two levels up, to the roughly 1.3 billion dollar programme the assistant is a deliverable of, and it arrives as budget lines and schedule discipline — never as a look at whether an answer was right. The figures that travel upward are the ones the assistant produces about itself, and what they count is how many conversations ended without a person. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. The Lab offers this deployment every tool its own record supports, and applying all of them at once costs two to three times the budget and still leaves pathways open, so no amount of money reaches the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the EDD-class two-tier benefits assistant with inherited oversight network: 13 components and 24 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 14 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the two-tier public benefits assistant documented in the California EDD case file — not a reconstruction of the actual service. The platform is documented as intent-based conversational AI; the public sources do not establish generative language modeling, and nothing on this diagram treats it as such.
- baseline
Binding framing ruling, applied structurally rather than only in prose: inherited oversight never touches this assistant's output. No pathway runs from either model to either reviewer. Both reviewers are wired in through the published service figures, because that is what the record describes — the department publishes usage counts that circulate into analyst documents and hearing agendas, so the metrics the assistant generates are themselves the evidence base the oversight layer reads. Nothing in that record engages with answer accuracy, translation fidelity, or the harm of a wrongly deflected contact.
- baseline
Demand reads 3 from documented load, not from a default: 554,792 unique customers and 2,103,782 messages on the public tier in the first half of 2025 alone, monthly users moving from 54,000 to 111,000 after migration, three benefit programmes paying roughly 2 million workers a year, and a programme created in response to a caseload the department could not absorb — 1.9 million rejected claims between March 2020 and October 2023 and a backlog of more than 130,000 pending appeals at a 137-day average wait.
- baseline
Capacity reads 1 because the human channel is bounded by design and the bound is documented precisely. Live agent chat runs weekdays from 9 a.m. to 2 p.m., 25 hours a week, and reached nearly 6,000 unemployment customers a month against roughly 111,000 monthly assistant users; the agency reports 29,900-plus customers served by live-chat agents in total. Deflection is the stated design goal, with the agency reporting 47 percent fewer customers needing an agent after self-service. Outside the weekday window the assistant is the only chat channel, which is why the override question in this deployment is whether a person can be reached rather than whether a person can overrule the machine.
- baseline
Two model components are drawn because the verified record keeps two systems apart and its cautions warn specifically against conflating them: an unauthenticated public-site assistant answering general programme questions in eight languages around the clock, and an authenticated chatbot launched on 8 May 2026 that reads a signed-in customer's own claim status, payment and eligibility for claims filed in the past three years. They have separate launches, separate scopes and separate usage counts, and merging them would merge counts that belong to different tiers.
- baseline
The published service figures are drawn as a record store rather than as prose because the record makes them structural: the department publishes usage counts, message totals, self-service action counts and resolution rates, programme leadership assembles and reads them, and they travel upward into oversight documents. The store's inbound and outbound edges are the highest in this network for that reason. What those figures count is channel exit — an 80-plus percent chat resolution rate and 47 percent fewer agent contacts measure conversations that ended in the channel, which is not the same as problems that were solved. They are treated here as an incentive-loop parameter, never as an outcome measure.
- baseline
The store-to-store pathway from the processing core into the portal account, and its reconciliation check drawn at zero, rest on a labeled inference and are marked as such. The oversight handout describes new front-end functionality as linked to the legacy core through informal and untested data bridges and custom-built interfaces, and says the functionality added since the pandemic has not been stress tested — but that handout does not name this assistant. The inference drawn here is that the authenticated tier, which reads individual claim data, rides those interfaces. The check is at zero because no reconciliation step appears anywhere in the sources; the core replacement is scheduled to complete in 2031.
- baseline
The second read of the assistant's answers is drawn at zero because its absence is documented, not because the shape looked bare. No independent audit of this service's answer accuracy exists in the public record: no inspector-general review, no state-auditor evaluation, no academic study, and no engagement with answer quality anywhere in the legislative or project-oversight record. Every performance figure that does exist is agency self-reported. Drawing this pathway lets the missing read be seen and closed rather than leaving it off the page.
- baseline
The identity check is drawn as a bounded automated screen carrying no flow of its own, and deliberately not as an output check on the model-to-operator pathway. What identity proofing and the three-year lookback bound is who may be answered and which claims are in scope — real, automated and partial by construction. Neither screens whether an answer is correct. Drawing an automated output check here would assert a control the record does not document, in a case whose central finding is that no such control appears anywhere.
- baseline
No enforcement node is drawn. The assistant is informational and never adjudicative: it does not determine eligibility or benefits, so there is no downstream action system and no algorithmic determination for a caseworker to overrule. Eligibility determinations stay with department staff through separate questionnaires and interviews, which are outside this diagram. No input-source node is drawn either: the platform is intent-based rather than retrieval-based, and the content the assistant answers from is maintained by the customer-experience function, which is drawn as the operator-to-model pathway the record supports.
- baseline
The queue is drawn as a mediator carrying no flow of its own because it is a documented quantity of this deployment rather than a pathway: more than 830,000 customers have used callback since May 2024, and the weekday 9 a.m. to 2 p.m. window is the interval inside which a live agent can pick up at all.
- baseline
The single-assistant self-loop is a language-access fact in this deployment rather than an abstraction. One assistant answers the whole state in California's top eight working-age languages and is the primary channel for claimants who do not use English, with six further languages machine-translated in the human chat, so a gap in one language's coverage repeats for every speaker of that language at once.
- baseline
The customer-experience pathway is the corrective loop this deployment actually has, and it is drawn at the middle rung rather than higher for a documented reason: its instruments are post-chat surveys, saved redacted transcripts and 4,500 hours of customer research, which measure satisfaction and experience. Those are real, and they are a proxy for answer correctness rather than a measure of it.
- assumed
The rungs on the two escalation pathways come from documented volumes rather than from a convention: the public tier's handoff runs at the middle rung because roughly 6,000 unemployment customers a month reach an agent against roughly 111,000 monthly assistant users, and the per-contact coupling is strong because the whole conversation, the verified identity and the account details travel with the customer. The authenticated tier's handoff runs one rung lower because the agency reports nearly 18,000 of more than 25,000 first-fortnight uses finishing in self-service, and because that tier was weeks old when this record was taken. Where the sources are silent — how often an agent's correction reaches the assistant's content, for instance — no pathway is drawn rather than a conservative one invented.
- baseline
The egress pathway to the contracted platform is drawn present rather than absent, and at the low rung rather than open. It is present because conversation content, and on the authenticated tier the claim details inside it, are genuinely processed on vendor infrastructure, which also performs the real-time machine translation of a human chat into six non-English languages. It is low because the crossing is governed: state contracts inside a programme with quarterly progress reporting and an expenditure-plan release gate.
- baseline
The vendor's own published figures — calls deflected per day, constituent hours saved, annual dollar savings — are integrator marketing claims with no independent verification, and no parameter on this diagram is derived from them. Separately, an agency-published cumulative user figure for 2025 that is hard to reconcile against the department's own precise unique-customer count for the same period is not used anywhere here.
- assumed
Served claimants are not in the dynamics. The people asking these questions, the answers they receive, the benefits they are or are not paid, and what happens to someone whose contact ends in the channel without their problem solved are boundary quantities recorded in the case file. The operator network — contact-centre agents, the service-design function and programme leadership — is what this diagram propagates, and no claimant outcome or differential harm is computed from anything drawn here.
What this example does not show
- BINDING FRAMING: the oversight in this network never touches the assistant's output, and the diagram is drawn that way on purpose. Analyst-office analyses, budget subcommittee hearings, project oversight and quarterly finance reporting attach to the modernization programme's budget, schedule and procurement. None of them engages with answer accuracy, translation fidelity, or the harm of a contact wrongly resolved in the channel, so no pathway on this diagram runs from either assistant to either review body.
- Every performance figure here is agency self-reported. Unique customers, message counts, self-service actions, the share of chats reported as resolved and the share of agent contacts avoided are published by the department in its own updates and blog posts. No inspector-general review, state-auditor evaluation, or academic study of this assistant's answer accuracy exists in the public record. The separate figures for calls deflected and dollars saved were published by the integrator on its own marketing blog and have never been independently verified; nothing on this diagram is derived from them. An agency-published cumulative user figure for 2025 is hard to reconcile against the department's own precise unique-customer count for the same period and is not used here.
- The link between the oversight record's data-bridge finding and this assistant is a LABELED INFERENCE. The handout attributes that risk generically to new front-end functionality reaching the legacy core through informal and untested data bridges and custom-built interfaces; it does not name the chatbot. What is drawn here is that the authenticated tier, which reads individual claim data, plausibly rides those interfaces. The completeness of the chat deliverables rests instead on the Senate subcommittee agenda's own plan lines.
- Two separate systems are kept separate here, and their counts never merge: an unauthenticated public-site assistant available since 2025 in eight languages, and an authenticated claim-information chatbot launched on 8 May 2026. The platform is documented as intent-based conversational AI; the public sources do not establish generative language modeling, and nothing here describes it as such.
- The success measures in this case are deflection-shaped. An 80-plus percent chat resolution rate and 47 percent fewer customers needing an agent measure channel exit, not verified problem resolution. They are read on this diagram as an incentive-loop parameter and never as an outcome measure.
- Served claimants are not modeled here. The people asking these questions, the answers they receive, the languages they ask in, the benefits they are or are not paid, and what happens to someone whose contact ends without their problem solved are boundary quantities recorded in the case file. The Lab models institutional propagation through the operator network, estimates no differential harm to served people, and computes no claimant outcome from anything on this diagram.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
California's Employment Development Department runs a two-tier chat assistant delivered under the Integrated Contact Center work stream of EDDNext, the state's roughly $1.258 billion modernization of its unemployment, disability, and paid family leave systems. The unauthenticated public-site tier became available around the clock in the state's top eight working-age languages in May 2025 and served 554,792 unique customers across 2,103,782 messages between January 1 and June 30, 2025. A live agent chat channel for unemployment customers, which the department dates to July 2025, lets a customer escalate to a person on weekdays between 9 a.m. and 2 p.m. after identity verification, with account details passed to the agent, real-time machine translation in six non-English languages, a redacted transcript saved to the account, and a post-chat survey. On May 8, 2026 a second, authenticated tier launched inside the customer portal, answering a signed-in unemployment customer's own claim status, payment, and eligibility questions for claims filed in the past three years; the department reported more than 25,000 uses and nearly 18,000 fully self-service interactions in its first two weeks. The platform is documented as intent-based conversational AI; the public sources do not establish generative language modeling. All usage figures are agency self-reported, and the two tiers' counts belong to different systems.
empirical- Government California Employment Development Department, Amazon Web Services Collaborates with the EDD to Improve Customer Service for Californians (2023) https://edd.ca.gov/en/newsroom/benefitting-californians/2023/amazon-web-services-collaborates-with-the-edd-to-improve-customer-service-for-californians/
- Government California Employment Development Department, EDD's Virtual Assistant (Chatbot) Now Available in Top Eight Languages (2025) https://edd.ca.gov/en/newsroom/benefitting-californians/2025/edds-virtual-assistant-chatbot-now-available-in-top-eight-languages/
- Government California Employment Development Department, EDD Modernization Update, July 30, 2025: Recent Customer Service Improvements (2025) https://edd.ca.gov/siteassets/files/pdf_pub_ctr/edd-modernization-update_july-2025.pdf
- Government California Employment Development Department, Unemployment Customers Can Chat Online with a Live Agent (2025) https://edd.ca.gov/en/newsroom/benefitting-californians/2025/unemployment-customers-can-chat-online-with-a-live-agent/
- Government California Employment Development Department, Unemployment Customers Can Now Get Claim Information Through Chat (2026) https://edd.ca.gov/en/newsroom/benefitting-californians/benefiting-californians-2026/unemployment-customers-can-now-get-claim-information-through-chat/
No body oversees the EDD chat assistant as such: its accountability is inherited two levels up, from legislative and executive scrutiny of the EDDNext programme it is a deliverable of. That scrutiny is substantial and has demonstrably changed programme behaviour — the 2026-27 Governor's Budget reverts $70.6 million of unused modernization funding early as a budget solution, and the core claims-system replacement was resequenced to do disability and paid family leave first with unemployment integration designated 'mandatory optional' to reduce risk for the state — and it is also documented as partial, since the adopted 2025-26 Budget Act retained extended spending authority against the Legislative Analyst's Office recommendation to drop it. What the oversight record engages with is budget, schedule, and procurement: the 2026 analyst-office questions on the record concern the core project's vendor and the rationale for removing unemployment from it, and the February 2026 handout flags that new front-end functionality is linked to the legacy claims mainframe through 'informal and untested data bridges and custom-built interfaces' that have not been stress tested — a finding stated generically about new functionality, without naming the chat assistant. No inspector-general review, state-auditor evaluation, or academic study of this assistant's answer accuracy or translation fidelity has been located in the public record.
empirical- Government California Legislative Analyst's Office, The 2025-26 Budget: EDDNext (2025) https://lao.ca.gov/Publications/Report/4985
- Government California Legislative Analyst's Office, The 2025-26 California Spending Plan: Other Provisions (2025) https://lao.ca.gov/Publications/Report/5081
- Government California Legislative Analyst's Office, Overview of Efforts to Modernize EDD's Benefit Systems (Assembly Budget Subcommittee No. 5 handout) (2026) https://lao.ca.gov/handouts/revtax/2026/Overview-of-Efforts-to-Modernize-EDD-Systems-022426.pdf
- Government California Legislative Analyst's Office, The 2026-27 Budget: Overview of the Governor's Budget (2026) https://lao.ca.gov/Publications/Report/5101
- Government California State Senate, Senate Budget and Fiscal Review Subcommittee No. 5 Hearing Agenda, April 23, 2026 (Issue 1: EDDNext Modernization) (2026) https://sbud.senate.ca.gov/system/files/2026-04/sub-5-4.23.26-hearing-agenda-final.pdf
The success measures published for the EDD chat assistant are deflection-shaped, and the department both produces them and reports them upward. Since July 2025 the agency states that 23 percent more customers resolve questions through self-service and 47 percent fewer need to speak with an agent after using it, alongside more than 2.1 million self-service actions since November 2024, more than 830,000 customers using callback since May 2024, more than 29,900 customers served by live-chat agents, and an in-chat issue-resolution rate above 80 percent. Those figures measure channel exit rather than verified problem resolution, and they are agency self-reported; the human channel they are measured against runs weekdays from 9 a.m. to 2 p.m. and reached nearly 6,000 unemployment customers a month as of October 2025. Separately, the integrator published on its own marketing blog that the bot stack reduced call volume by 35,000 calls a day, cut live-agent volume by 3,800 calls a day, saved constituents 684 hours daily, and generated $8.4 million in annual savings; those are vendor claims with no independent verification.
empirical- Government California Employment Development Department, Smarter Service: Inside EDD's Contact Center Modernization (2025) https://edd.ca.gov/en/newsroom/benefitting-californians/2025/smarter-service-inside-edds-contact-center-modernization/
- Government California Employment Development Department, Unemployment Customers Can Chat Online with a Live Agent (2025) https://edd.ca.gov/en/newsroom/benefitting-californians/2025/unemployment-customers-can-chat-online-with-a-live-agent/
- Vendor InterVision Systems, InterVision's Contact Center Services Have Transformed California's EDD (vendor blog) (2025) https://intervision.com/blog-contact-center-edd/
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Keep prompts neutral — Framing and mirroring reduction
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Vet connections — Connection authorization
- Check copied records — Reconcile copied records
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Gate vendor updates — Vendor quality gate
Documented case histories
- EDD Virtual Assistant
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down