PAN Lab example
CDTFA Axyom Assist
The renewal decision: a call-center answer assistant
A state tax department buys a retrieval assistant for its call center. It listens to the live call, finds candidate answers in 16,000-plus pages of the department's own published material, and drafts the summary afterwards. An agent reviews everything; nothing generated reaches the caller unreviewed. Modeled on the California tax department's call-center deployment. What makes this network unusual is where the strongest checking sits: not on the desk, but on the contract. Two vendors were tested for six months at a dollar each, one was cut, the survivor ran departmentwide for a year, and then the purchaser read the benefit case against the price and let the contract lapse — substituting an off-the-shelf capability under a contract it already had. The saving that justified the build was a projection from a simulated pilot, and it never became a production measurement. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. With every tool the Lab currently offers, no affordable combination brings this system inside the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the CDTFA-class governed-procurement call-center assistant network: 11 components and 22 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 10 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the governed procure-evaluate-substitute pattern documented in the California CDTFA call-center assistant case file — not a reconstruction of the actual tool.
- baseline
Demand reads 3 from documented load, not from a default: more than 800,000 contacts a year across phone, chat and email against roughly 375 agents, peak filing season at 10,000 calls a day (about four times normal), waits moving from about 4 to about 20 minutes, and roughly 280 staff historically reassigned out of other units to absorb the peak.
- baseline
Capacity reads 3 because the human process here worked. Agents answered these contacts by searching the same reference corpus by hand, the project's own stated goal was to shorten that research time, the pilot projected only a minimum 1.5 percent per-call saving on top of it, and when the contract ended the department moved to an off-the-shelf capability under an existing contract. That high counterfactual floor is the reason the corpus-to-agents pathway is drawn at full strength.
- baseline
The benefit arithmetic is a projection, not a measurement. The call-center chief described a simulated environment that showed potential to save at least one and a half percent of time on some calls; the roughly 100,000 minutes a year and about 10,000 additional calls are agency extrapolations from that figure across about 800,000 annual contacts, relayed through trade press and echoed in vendor marketing. No production-environment measurement, methodology or assessment report has been located, which is why the model-side cross-check pathway is drawn at the low rung and the human-centered assessment pathway with it.
- baseline
The store-to-store pathway from summaries into the knowledge corpus is drawn at zero because the record documents its absence, not because the shape looked bare: the corpus was limited to publicly available department material by the safeguards stated at launch, and it was not fed by the assistant's own outputs. The one residual write channel is a drafted post-call summary entering internal workflow records after agent review; the assistant has no write path into a taxpayer's case record.
- baseline
The retrieval layer and the contact queue are drawn as mediators carrying no flow of their own. Retrieval is a named component of this deployment rather than a figure of speech — the product retrieves candidate passages from 42 programs before generating — and the queue is the quantity the entire procurement judgment was about, with published waits of about 4 minutes off-season against about 20 at peak.
- baseline
Two reviewer nodes are drawn because the record documents two bodies that reached different conclusions: an interagency evaluation layer (operations agency, state technology department running the isolated sandbox, general-services department, and the data-innovation office charged with the human-centered final assessment) and a purchaser-side renewal authority. Both are wired in with their own inbound pathway from the assistant. The renewal authority carries the strongest inhibiting edge in this network, because in this deployment it is the pathway that demonstrably fired: once cutting the vendor field from two to one, once ending a working contract.
- baseline
No enforcement node is drawn. The record documents no scored subject, no adverse-action channel and no autonomous action path — the assistant never interacts with the taxpayer. No automated output screen is drawn either: the vendor's contractual duty on the record is to monitor and report on response accuracy, coherence and appropriateness, which is a reporting obligation into the evaluation layer rather than an inline gate before an agent sees a suggestion. The documented pre-delivery control is mandatory human review, carried by the bounded model-to-operator rungs and the operator check pathways.
- assumed
The two operator classes differ by corpus fluency, not by review regime. Both receive the same suggestion stream under the same mandatory-review requirement, so their model-to-operator rungs match; the documented difference is that roughly 280 staff are moved in from other units for the filing peak and read a 42-program corpus they have not spent the year inside, which is where an outside consumer group's launch-time question about whether call-center staff can verify generated answers lands. Peer spread and coaching between the classes follow from the documented co-working at peak; the record gives no rates for either, and the rungs here are conservative.
- assumed
The follow-up read of a prior summary is drawn at the low rung because the published record is silent on how often a previous call's summary is consulted on a repeat contact. Where the record is silent this network takes the conservative value rather than importing another deployment's number.
- baseline
The egress pathway to the vendor-hosted service is drawn present rather than absent, and at the low rung rather than open. It is present because live taxpayer conversation content genuinely crosses to a contracted vendor environment for transcription and generation — the publicly-available-data-only safeguard governed the knowledge corpus, not the call. It is low because the crossing was governed: a state contract, a six-month isolated sandbox operated by the state technology department during evaluation, and state oversight of that evaluation.
- baseline
One assistant on one corpus served every desk in the department once implementation completed in August 2025, so a gap in its answers repeats across the whole floor rather than case by case; the model self-loop encodes that correlated reach.
- baseline
A separate department proof-of-concept, reported as hindered by human error, is not this deployment. That error was a state technology staffer's miscalculation in the vendor-evaluation scoring used during procurement, on a distinct machine-learning consulting contract. Nothing about it is attributed to this assistant, its data, or its corpus, and no corpus-quality feedback pathway is drawn from it.
- assumed
Served taxpayers are not in the dynamics. Callers, their questions, the answers they receive and the time they spend waiting are boundary quantities recorded in the case file; the operator network is what this diagram propagates. The assistant never interacted with the taxpayer by design, and no caller outcome or differential harm is computed from anything drawn here.
What this example does not show
- The benefit figures here are PROJECTIONS from a simulated pilot environment, self-reported by the agency through trade press and echoed in vendor marketing — a minimum of about 1.5 percent per call, extrapolated to roughly 100,000 minutes a year and about 10,000 additional calls. No production-environment measurement, methodology, or independent assessment report has been located, and nothing on this diagram measures the saving.
- The end of this contract was a procurement authority action, not an adjudicated failure. The department declined to renew and moved to a similar off-the-shelf capability under a contract it already held: governed substitution under commoditization pressure, not abandonment of the technology and not a finding by any tribunal that the system failed. The marginal cost of the replacement is not public, so the cost side of that judgment is quoted rather than itemized.
- A separate department proof-of-concept reported as hindered by human error is a different project. That error was a state technology staffer's miscalculation in the vendor-evaluation scoring used during procurement, on a distinct machine-learning consulting contract, and nothing about it is attributed to this assistant, its data, or its corpus.
- Served taxpayers are not modeled here. Callers, the questions they ask, the answers they receive and the time they spend waiting are boundary quantities recorded in the case file; the Lab models institutional propagation through the operator network, estimates no differential harm to served people, and computes no caller outcome from anything on this diagram.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
California's tax agency ran the first state generative AI call-center assistant through a complete governed procurement arc: a competitive sandbox in which two vendors were each paid one dollar to test for six months in a secure environment, a 10-month pilot in a simulated environment, and production rollout to roughly 375 agents completed August 22, 2025 under a 12-month, 445,000 dollar contract. The pilot projected a minimum 1.5 percent per-call time saving, and by May 2026 the agency had declined to renew the contract, with the official who launched the project saying the working system 'didn't quite save as much time as we had hoped.'
empirical- Government State of California GenAI portal (genai.ca.gov), California signs partnerships to utilize GenAI (2024) https://www.genai.ca.gov/2024/05/09/california-signs-partnerships-to-utilize-genai/
- Government Office of the Governor of California, Governor Newsom deploys first-in-the-nation GenAI technologies to improve efficiency in state government (2025) https://www.gov.ca.gov/2025/04/29/governor-newsom-deploys-first-in-the-nation-genai-technologies-to-improve-efficiency-in-state-government/
- Trade press Government Technology Industry Insider California, CDTFA Moves GenAI Assistant Into Call Center Production Environment (2025) https://insider.govtech.com/california/news/cdtfa-moves-genai-assistant-into-call-center-production-environment
- Investigative Melhado, CA agencies discontinue some AI projects aimed at making government more efficient (The Sacramento Bee via Yahoo News) (2026) https://www.yahoo.com/news/articles/ca-agencies-discontinue-ai-projects-120000898.html
The benefit figures for the CDTFA call-center assistant are pilot projections from a simulated environment, self-reported by the agency through trade press and echoed in vendor marketing: the call-center chief said the simulated environment 'did show some potential to save at least one and a half percent time on some of those calls,' extrapolated to roughly 100,000 minutes per year and capacity for about 10,000 additional calls annually across roughly 800,000 yearly inquiries. No published production-environment measurement, methodology, or independent assessment report has been located.
empirical- Trade press Government Technology Industry Insider California, CDTFA Moves GenAI Assistant Into Call Center Production Environment (2025) https://insider.govtech.com/california/news/cdtfa-moves-genai-assistant-into-call-center-production-environment
- Vendor SymSoft Solutions via Business Wire, SymSoft Solutions Powers California's GenAI Revolution with Axyom Assist at CDTFA (press release) (2025) https://www.businesswire.com/news/home/20250514665066/en/SymSoft-Solutions-Powers-Californias-GenAI-Revolution-with-Axyom-Assist-at-CDTFA
By May 2026 CDTFA had declined to renew the 12-month, 445,000 dollar contract for its custom call-center assistant and shifted to a similar off-the-shelf tool from Amazon Web Services under an existing larger contract, with Government Operations Secretary Nick Maduros saying the solutions 'worked in practice' but 'didn't quite save as much time as we had hoped.' The decision was a procurement authority action inside a broader portfolio review of eight executive-order-driven generative AI pilots totaling more than 5.8 million dollars, in which three other projects also ended while three continued.
empirical- Investigative Melhado, CA agencies discontinue some AI projects aimed at making government more efficient (The Sacramento Bee via Yahoo News) (2026) https://www.yahoo.com/news/articles/ca-agencies-discontinue-ai-projects-120000898.html
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Keep prompts neutral — Framing and mirroring reduction
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Store less data — Data minimization
- Keep skills sharp — Deskilling-arrest mandate
- Peer sharing rules — Peer-edge governance
- Check with a second model — Cross-model verification
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Gate vendor updates — Vendor quality gate
Documented case histories
- CDTFA Axyom Assist
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check