PAN Lab example
Mass.gov Virtual Assistant
The pages it reads, rewritten so it reads them better
A state built its own assistant and put it in front of the safety net. It answers questions about food assistance, unemployment, child support and family leave from a knowledge base rebuilt every night out of the state's own published pages. Modeled on a real state deployment - its shape, not the tool itself. There is no caseworker in the middle and no case file at the end, so a wrong answer leaves no appeal trail and the only thing that can see it is an evaluation tool the same team built. Then the interesting part: the pages were updated and simplified so the assistant would answer from them well. The authoritative public statement of what a benefit rule says is being written for the machine that reads it. Two reviewers exist. Both read the catalog entry rather than the answers, one has an assessment instrument whose completions the released entries record at zero, and the one that changed the agency's behavior came from records law and reached the paperwork. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. The Lab offers this deployment every tool its own record supports, and applying all of them at once costs two to three times the budget and still leaves pathways open, so no amount of money reaches the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Multi-agency public assistant over co-authored program pages network: 11 components and 17 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 7 assumed · 4 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the multi-agency public-assistant pattern documented in the massgov-virtual-assistant case file — a state-built retrieval-grounded chatbot standing in front of safety-net program content, whose ground truth its own operators write and whose only view of its own errors is tooling it built itself — and not a reconstruction of the actual assistant. The foundation model and the hosting stack are undisclosed in every source located and are named nowhere in this network.
- assumed
Every performance figure in this record is agency self-reported and unaudited, and the records that would permit independent verification were withheld. More than 1,500 conversations a day, roughly 200,000 conversations since launch, a chat open rate rising from 1.29 to 2.92 percent, positive feedback rising from 10 to 56 percent, negative feedback down 60 percent and page abandonment down 40 percent all come from the deploying agency; the cost and usage reports and technical logs that would confirm its claims were withheld from a records request. No edge weight, demand value or capacity value here is calibrated to any of those numbers.
- baseline
The assessment finding is drawn at exactly its verified scope, per the binding framing ruling. The documented fact is that none of the nine inventory entries released to a records request recorded a completed privacy impact assessment, with the fields left blank and no explanation offered; the other thirty-one entries were withheld, so a count across all forty is an inference the agency has not contradicted rather than a verified number. The assessment is an internal enterprise-privacy-office practice described in a report to the legislature, and the policy in force names no assessment requirement of its own, so the latent check on this diagram is an internal commitment and never a violated legal requirement.
- baseline
The reviewer wiring follows the evidence rather than the corpus habit. Both reviewers read the same object — the generative-AI use-case inventory — because that is what the record documents each of them reaching: the privacy office through its assessment fields, the records supervisor through a public-records request that released nine entries and was refused thirty-one. Neither reviewer receives the assistant's answers, and that is the finding, not an omission: the oversight in this deployment attaches to the documentation about the system and never to what the system says.
- baseline
The source-record co-adaptation is drawn as this network's load-bearing structure because three documented facts compose into one loop: the published program pages are the state's authoritative statement of the program rules, those pages were updated and simplified so the assistant would answer from them well, and the knowledge base is rebuilt out of them every night. Two operator classes write into that store and the assistant reads it, so the ground truth and the channel that reads it move together. The protective and the cautionary readings come from the same fact: simpler public pages are a real gain for anyone reading them, and a public record shaped for a machine channel is a record whose audience has quietly changed.
- assumed
A heavy workload against limited capacity is derived, not defaulted. Demand is high on documented volume and scope: more than 1,500 conversations a day around the clock across seven program areas and the state login service on one platform, a build run in 83 days against a statutory identification deadline, an explicit official framing of AI as closing a gap for stretched agencies without adding headcount, and a union representing about 9,000 state employees arguing that staffing rather than AI is the capacity constraint. Capacity is the competent-baseline value because the human channel worked and still runs: the reported reduction of about 1,000 calls a day and 200 emails a week is a measure of what people were doing before, and the predecessor rule-based chatbot covered twenty languages against this deployment's three.
- assumed
The pathway from the published pages to the call-center desk is a labeled inference and is modeled as one. The sources establish those pages as the state's authoritative public statement of these program rules and document the phone channel's contact volume falling after launch; they do not describe how the desk reads the pages. It is carried below the top step for that reason, and no figure in this network is derived from it. The same caution applies to the phone class generally: what the record measures about it is contact volume, not skill or capability.
- assumed
The bounded evaluation screen is drawn at a low level and no higher. It is drawn at all because the deploying agency describes a custom tool that flags answer issues within hours for staff correction; it is held to a single step because such a screen is partial by construction, because it acts once an answer has already reached the person who asked, and because the technical logs and usage reports that would show what it catches were withheld from a records request. Its placement on the platform team's channel follows the record: that team is the only one with any view of the assistant's output.
- baseline
Two pathways are drawn without flow because sources document them as absent rather than merely unobserved: the privacy assessment of a catalogued use case, and an independent read of what the assistant gets right and wrong. Each is drawn so the gap is visible on the diagram and a lever can reach it. Drawing them is not a claim that either was legally required; the assessment is an internal commitment, and no independent evaluation of this assistant appears in any source located.
- assumed
Three adjacent parts of the state's portfolio are deliberately absent from this network. The call-summarization pilot that writes machine-made summaries into a benefits record is the separate copilots case and may never be attributed here. The vendor-hosted components of the wider portfolio, including a third-party-liability tool held on a vendor's servers, belong to other use cases and not to this assistant, which the state describes as a state-owned platform. And the executive-branch-wide workforce assistant procured in February 2026 is a different deployment whose figures appear nowhere in this network. Scope here runs to the program content the assistant's own support page lists and not to the state health insurance program, which that page does not list.
- assumed
Served residents are not in the dynamics. The people typing into this assistant are members of the public applying for or asking about benefits, and they are the direct users of the model, which is precisely why nothing about their outcomes is computed here: the Lab models institutional propagation only. Two differential observations sit in the case file with their labels and are never derived from this diagram — the assistant offers three languages where the predecessor rule-based chatbot offered twenty, and the human fallback channel is narrowing on the agency's own success figures. Neither effect on any resident group is measured in any source located, and an error here produces no case record, no adverse-action notice and no appeal trail through which it could be counted.
What this example does not show
- Served residents are not modeled here. The people typing into this assistant are members of the public asking about their benefits, and they are the direct users of the model - which is exactly why nothing about their outcomes is computed on this diagram. The Lab models institutional propagation only; an error here produces no case record, no adverse-action notice and no appeal trail, and the case file carries what the record says about who is affected.
- The assessment finding is stated at its verified scope: none of the NINE inventory entries released to a records request recorded a completed privacy impact assessment, and the other thirty-one were withheld, so a count across all forty is an inference the agency has not contradicted rather than a verified number. The assessment is an internal enterprise-privacy-office practice described in a report to the legislature, and the policy in force names no assessment requirement of its own, so this is an unexercised internal commitment and never a violated legal requirement.
- Every performance figure in this record is an agency self-report and unaudited: the conversation counts, the chat open rate, the feedback shifts, the abandonment figure and the call and email reductions all come from the deploying state or from an outside brief reporting the state's own numbers, and the cost and usage reports and technical logs that would permit independent verification were withheld from a records request. No independent evaluation of this assistant's accuracy appears in any source located.
- The assistant is state-built on what the agency calls a state-owned platform, and its underlying foundation model and hosting stack are undisclosed and are named nowhere in this scenario. Vendor hosting in this state's wider AI portfolio belongs to other use cases and is never asserted of this assistant. The call-summarization pilot that writes machine-made summaries into a benefits record is a separate case and is not attributed here, and scope runs to the program content the assistant's own support page lists rather than to the state health insurance program, which that page does not list.
- The order to produce the withheld records for in camera review is a procedural step and not a finding: no final determination appears in the record as checked to July 2026, and the month of the order is an inference from the phrase used in the reporting rather than a documented date. The union figure is about 9,000 state employees represented, not total membership.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Massachusetts built a retrieval-grounded generative-AI Virtual Assistant in house in 83 days and launched it on motor-vehicle registry pages in April 2025 ahead of the May 7, 2025 federal identification deadline, describing it as a state-owned platform that replaced a vendor-managed rule-based chatbot; by 2026 the assistant's own support page listed motor-vehicle, toll, unemployment-assistance, tax, child-support, transitional-assistance, family-and-medical-leave, and state-login content. The state reports more than 1,500 conversations a day and roughly 200,000 conversations since launch, a chat open rate rising from 1.29 to 2.92 percent, positive feedback rising from 10 to 56 percent, negative feedback down 60 percent, page abandonment down 40 percent, around-the-clock availability in English, Spanish, and Portuguese, a knowledge base refreshed nightly from published state content, and a custom evaluation tool that flags answer issues within hours. Every one of those figures is an agency self-report; the technology agency withheld the cost and usage reports and the technical logs that would confirm its claims, and the underlying foundation model and hosting stack have not been publicly identified.
empirical- Government Massachusetts Digital Service, Launching the Commonwealth's first generative AI Virtual Assistant (2025) https://www.mass.gov/info-details/launching-the-commonwealths-first-generative-ai-virtual-assistant
- Government Executive Office of Technology Services and Security, Mass.gov Virtual Assistant Inquiry/Support (2026) https://www.mass.gov/how-to/massgov-virtual-assistant-inquirysupport
- Government Executive Office of Technology Services and Security, Annual Legislative Report pursuant to Chapter 140 of the Acts of 2024, filed as HD4511 (2025) https://malegislature.gov/Bills/194/HD4511.pdf
- Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
- Government Massachusetts Digital Service, Delivering on the Digital Roadmap (2026) https://www.mass.gov/info-details/delivering-on-the-digital-roadmap
Massachusetts policy AI.001, effective January 31, 2025, requires human fact-checking of generative output, conspicuous labeling of AI content, Chief Technology Officer approval for generative procurement, and a generative-AI inventory, while routing privacy review through consultation with legal and security teams; the phrase privacy impact appears in it zero times. The Enterprise Privacy Office told the legislature in February 2025 that it has been using a Privacy Impact Assessment and was overseeing a pilot program with its risk and security teams to assess privacy risks during the contracting and development stages. Records obtained by an independent investigation showed at least forty AI use cases in the agency's internal survey with thirty-one withheld, and of the nine entries released not one recorded a completed privacy impact assessment, the fields left blank without explanation including for tools processing Social Security numbers and Medicaid data, with the agency spokesperson declining to answer questions about them; a count across all forty is therefore an inference the agency has not contradicted rather than a verified number, and the assessment is an internal office practice rather than a statutory mandate. After months of negotiation the Supervisor of Records ordered the agency to produce the withheld records for in camera review, allowing ten business days to produce and up to fifteen business days to review; no final determination was found as of July 20, 2026.
empirical- Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
- Government Commonwealth of Massachusetts Enterprise Privacy Office, Enterprise Use and Development of Generative Artificial Intelligence Policy AI.001, effective January 31, 2025 (2025) https://www.mass.gov/doc/enterprise-use-and-development-of-generative-artificial-intelligence-policy/download
- Government Executive Office of Technology Services and Security, Annual Legislative Report pursuant to Chapter 140 of the Acts of 2024, filed as HD4511 (2025) https://malegislature.gov/Bills/194/HD4511.pdf
- Government Massachusetts Executive Office of Technology Services and Security, Artificial Intelligence at the Commonwealth (Mass.gov) (2026) https://www.mass.gov/artificial-intelligence-at-the-commonwealth
The Massachusetts assistant answers from a knowledge base rebuilt nightly out of published state content, and the state reports that those published pages were updated and simplified to suit the assistant while chat analytics drive prompt tuning, content updates, and enhancements, so the authoritative public statement of program rules is adapted to the channel that reads it. On the human-channel side, an independent institute brief reporting state figures records motor-vehicle registry calls falling by about 1,000 a day and emails by about 200 a week after launch, counts about twenty AI use cases publicly reported with three facing the public against forty in the internal survey, and recommends a formal public AI inventory and structured user feedback loops; the assistant offers three languages where the predecessor rule-based chatbot offered twenty. The state technology secretary committed publicly that there will be a human reviewing output before public distribution and framed AI as relieving stretched agencies without adding headcount, while a public-employee union representing about 9,000 state employees argued that staffing rather than AI tooling is what would let the transitional-assistance agency serve more clients.
empirical- Government Massachusetts Digital Service, Launching the Commonwealth's first generative AI Virtual Assistant (2025) https://www.mass.gov/info-details/launching-the-commonwealths-first-generative-ai-virtual-assistant
- Investigative Pioneer Institute, Massachusetts has taken an important step on government AI but the Commonwealth must do more to improve services, transparency and save taxpayer dollars (Gary Blank) (2026) https://pioneerinstitute.org/massachusetts-has-taken-an-important-step-on-government-ai-but-the-commonwealth-must-do-more-to-improve-services-transparency-and-save-taxpayer-dollars/
- Investigative CommonWealth Beacon, What it means that a state AI assistant will handle your data, The Codcast interview with Secretary Jason Snyder (2026) https://commonwealthbeacon.org/the-codcast/what-it-means-that-a-state-ai-assistant-will-handle-your-data/
- Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
- Government Executive Office of Technology Services and Security, Annual Legislative Report pursuant to Chapter 140 of the Acts of 2024, filed as HD4511 (2025) https://malegislature.gov/Bills/194/HD4511.pdf
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Keep prompts neutral — Framing and mirroring reduction
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Peer sharing rules — Peer-edge governance
- Vet connections — Connection authorization
- Store less data — Data minimization
- Check with a second model — Cross-model verification
- Understand the system — Understand the system
- Upgrade model — Improve the model
Documented case histories
- Mass.gov Virtual Assistant
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down