PAN Lab example
Massachusetts DTA call summaries
The record and its source: AI summaries of benefits calls
A state benefits agency pilots a vendor-built copilot that transcribes eligibility calls in real time and writes a structured summary the caseworker can edit before saving it into the eligibility system of record. The full recording is then discarded by design; only the summary persists. Watch where the error travels. The model never speaks to the determination, recertification and hearings staff who act on its text - it reaches them through the record, as the account of an interview they were not present for, in a pipeline that dispositioned 31,390 applications in one month. And watch the one check that would catch a mistake: setting the saved summary beside the call it came from. The agency frames non-retention as a privacy protection, so on this board a privacy control and an accountability control sit on the same design choice. About 400 calls had gone through the tool by the April 2026 reporting, against 45,703 callers connected to staff in December 2025 alone - a pilot with room to expand. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. With every tool the Lab currently offers, no affordable combination brings this system inside the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Call-summary-class benefits documentation copilot network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 6 assumed · 14 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the call-documentation-copilot pattern documented in the Massachusetts Department of Transitional Assistance (DTA) call-summaries case file, not a reconstruction of the actual pipeline. The summarization prompt is withheld as proprietary, so the transformation applied between the call and the record is not public and is not reconstructed here.
- baseline
The reconciliation check points at the call audio rather than at another store, and that is the whole shape. The documented design keeps the caseworker-editable summary and discards the full transcript, so the step that would set the saved record beside the account it was derived from has nothing left to read. It is drawn dark for that documented reason, not because a review was skipped. This is the case's central tension in structural form: the agency frames non-retention as a privacy protection, so a privacy control and an accountability control sit on the same design choice.
- assumed
No pathway runs from the case record back into the model. The pipeline reads the live call, and the record documents no retrieval of prior case text into the summary and no retraining loop over the store; the deployment's own documentation listed its interaction-data plan as still to be defined while the tool was in production. Most documentation copilots in this catalogue do draw that retrieval pathway, and a few others do not; drawing it here would assert a reuse pathway the record does not describe.
- baseline
Two operator classes are drawn because the record describes two groups with different relations to the evidence: the worker who took the call, heard it, and may edit the summary, and the determination, recertification and hearings staff who later read that summary as the account of an interview they were not present for. The asymmetry between them is the case, not a stylistic choice.
- baseline
Inflow from the call audio is set at the maximum because the whole two-party call enters the pipeline unredacted, including Social Security numbers, medical history and immigration status, and no filter or redaction between the audio and the model appears anywhere in the record. The same reasoning marks that pathway, and the write of the finished summary into the case record, as privacy-sensitive.
- baseline
The two worker-side pathways are set from the documented gates rather than from volume. The structured summary reaches the worker who took the call as the draft standing in for the note they would otherwise have written themselves, which is the agency's stated purpose for the tool, so that pathway sits one step below the maximum: an edit step exists and worker use is voluntary under the collective agreement. The return pathway carries the two documented worker gates, which calls run through the tool and the edit before the save, and sits at the same level rather than higher, because the summarization prompt is the vendor's and is withheld, so the worker does not shape the transformation itself.
- baseline
The save into the case record is set at the maximum because every call that enters the pipeline ends there: on save the summary becomes the client's case record, as the account of the interview. The write-back from the determination, recertification and hearings staff sits one step below it, because what returns to the record is a decision line reached from the summary rather than the interview narrative itself.
- baseline
Two further pathways are drawn dark on the documented design rather than on missing information. A machine write reaching the case record without passing a worker's save is not part of the documented design, which routes every write through that save step; it is drawn because it is the pathway the expansion pressure this scenario carries would act on. The route by which a later reader puts a disputed line back to the worker who took the call is drawn for the same reason: the record states that after the save there is no override surface, and describes the correction burden falling instead on the client's memory against the state's system of record.
- baseline
The downstream read is set at the maximum on documented scale and documented dependence: the summary is the durable narrative of the interview, and future determinations, recertifications, discrepancy checks and fair hearings consult it, across 31,390 application dispositions in December 2025, 56,454 recertifications due, 24% monthly churn and an average of 12 days to approve a new application.
- baseline
The peer pathway between caseworkers is set low on coverage rather than on intensity. Use is voluntary under the collective agreement, so the tool travels by colleague example; about 400 calls had gone through it between the December 2025 rollout and the April 2026 reporting, against 45,703 callers connected to staff in December 2025 alone, which is well under 1% of connected calls. This is a pilot with room to expand rather than a saturated deployment, and the network is drawn that way.
- baseline
The model self-loop is set at moderate rather than maximum for the same coverage reason. One pipeline and one withheld prompt handle every call that enters the pilot, so a weakness in the transcription or the prompt repeats across the whole processed caseload instead of averaging out. The caller population is heavily multilingual, with Spanish, Haitian Creole, Chinese, Portuguese and Vietnamese the top languages after English, and includes 263,828 recipients aged 60 or over and 308,652 with a disability, so a uniform pipeline's weaknesses land on the same groups each time. That heterogeneity is a documented composition of the caseload, never a differential harm computed here.
- baseline
The governance node is the executive technology office that owns the state's generative-AI policy and internal AI use-case inventory. Its inbound pathway is a registration rather than a review feed: the deployment appears as one of the nine use cases released from an inventory of at least forty, and what travels to the channel is the fact of the tool rather than the summaries it writes. Its outbound assessment pathway is dark because none of the nine released entries, this one included, reported a completed privacy impact assessment even though the spreadsheet carried a field for it, and the tool's interaction-data plan was listed as something to be defined during requirements gathering while it was already in production.
- baseline
The union is deliberately not modeled as a component. The collective agreement, reached pre or early deployment, is documented as making worker use voluntary and adding job protections; it reviewed no output and did not address record provenance or client-side safeguards. Its documented effect is a worker-side selection gate over which calls enter the pipeline, which is why it is drawn as the worker-to-model pathway rather than as a reviewer. The record does not establish that the agreement was concluded before the December 2025 rollout, and nothing here assumes it was.
- assumed
The callers are not on the network and never enter the dynamics. The caller-side notice and opt-out before connection is an agency claim with no published opt-out rate, and the people whose calls are summarized are boundary-only, so that gate is recorded in the case file rather than drawn as an element.
- baseline
Exactly one data-leaving pathway is drawn, and it is dark. While the pilot ran, a federal agency demanded personal applicant and recipient data; preliminary injunctions of October 15, 2025 and February 27, 2026 in California v. USDA, No. 3:25-cv-06310 (N.D. Cal.) blocked funding cuts over the states' refusal, the court holding the proposed protocol would likely permit sharing beyond the entities allowed under 7 U.S.C. 2020(e)(8). Both orders are preliminary and the litigation is live, and whether AI summaries in the case record fall within the demand is an open question the state Attorney General's office declined to address. The other two boundary crossings are not drawn: no worker paste-out to an unsanctioned tool is documented for this deployment, and the record describes summary data as held in state-owned systems rather than on an ungoverned host, which is itself an agency claim and is labeled as one. The separate statewide enterprise AI contract is a different track and is not conflated with this pilot.
- assumed
No second-model check pathway is drawn. None is documented or planned, no independent evaluation of the tool has been published, and a second model would face the same discarded recording, so the model side carries only its own self-loop. An undocumented pathway is left off the board rather than added as decoration.
- baseline
Standing workload is set high and the no-AI counterfactual at a competent human baseline. The assistance line connected 45,703 callers to staff in December 2025, with daily averages of 2,709 calls connected, 5,427 completed in self-service and 6,454 callers unable to connect at all, for a program serving 1,011,460 recipients in 622,837 households, one in six state residents. The counterfactual is deliberately not set high: nothing documents automation displacing a demonstrably better working process, no evaluation exists in either direction, and the manual process also produced no verbatim record, since caseworkers wrote their notes from memory and no recording was kept then either. What the deployment changes is who authors the account, not whether a recording survives.
- assumed
The agency's benefit claims are unmeasured and are treated as claims. Reduced call handle time, improved consistency of case notes, freeing workers to focus on the conversation, storage in state-owned systems, compliance with existing access-control and retention policies, and caller opt-out availability are all agency statements; no independent evaluation, inspector-general audit or completed privacy impact assessment of this tool is on record, and edit rates and edit depth at the worker gate are unpublished. No benefit figure is derived on this diagram.
- baseline
The structural picture rests on single-outlet corroboration. The transcript non-retention finding, the call-count figure, the withheld prompt and the blank assessment fields all trace to one investigative outlet citing technical documentation and on-record agency statements. The caseload, call-volume, disposition, language, age and disability figures come separately from the agency's own published performance scorecard, which notes that its churn rate was recently updated following a computational error. No adjudicated individual harm exists; the concern this shape encodes is structural and forward-looking.
- assumed
Differential effects on the people who call are documented, never computed. The caseload is disproportionately elderly, disabled and non-English-speaking, and transcription quality across those groups has not been measured; because the recordings are discarded, the evidence base that would support such a measurement is discarded with them. That is recorded in the case file. This Lab models institutional propagation and estimates no harm to served people.
What this example does not show
- No outcome for any caller is modeled. The Lab reads institutional propagation only; Supplemental Nutrition Assistance Program (SNAP) applicants and recipients are boundary-only, and the caseload composition, the correction burden that falls on a client's memory against the state's system of record, and the caller-side notice and opt-out all live in the case file, never computed on this diagram.
- The harm this network encodes is structural and forward-looking, not adjudicated. No named claimant whose benefits were affected by a summary error appears in the record, and nothing here asserts one. The concern is that the material needed to establish such a case is discarded by design before the case could be brought.
- The structural picture rests on single-outlet corroboration. The transcript non-retention finding, the call count, the withheld prompt and the blank assessment fields all trace to one investigative outlet citing technical documentation and on-record agency statements; no inspector-general audit, agency evaluation or court filing about this specific tool exists. The caseload, call-volume, disposition, language, age and disability figures come separately from the agency's own performance scorecard, which notes that its churn rate was recently updated following a computational error.
- Every agency benefit and safeguard statement is an agency claim, labelled as one and never measured here: reduced call handle time, improved note consistency, storage in state-owned systems, compliance with existing access-control and retention policies, and the availability of a caller opt-out. Opt-out rates, summary edit rates and edit depth are unpublished.
- This is a pilot, and the board is drawn as one. Roughly 400 calls had gone through the tool by the April 2026 reporting, against roughly 45,000 callers connected to staff per month - well under 1% of connected calls. It is modeled as a deployment with room to expand, not as a saturated one.
- The litigation over the record store is live and its orders are preliminary. Preliminary injunctions issued on October 15, 2025 and February 27, 2026 in California v. USDA, No. 3:25-cv-06310 (N.D. Cal.); whether AI-generated summaries held in the eligibility record fall within the demanded data is an open question the state Attorney General's office declined to address, and it stays open here.
- The collective agreement's timing is hedged everywhere, including on the diagram: the record establishes that an agreement making worker use voluntary and protecting jobs exists, but not that it was concluded before the December 2025 rollout.
- This deployment is not the state's separate February 2026 enterprise AI contract. The documented legislator and union pressure targets that contract; any bearing on this pilot is inferential and is not drawn as a pathway here.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In December 2025 the Massachusetts Department of Transitional Assistance piloted an Accenture-built tool that transcribes SNAP eligibility calls in real time and generates a structured, caseworker-editable summary that is saved into BEACON, the state's benefits eligibility system of record; full transcripts are not retained, only the summaries, according to technical documentation reviewed by The Shoestring, and the summarization prompt was withheld as proprietary. About 400 calls had been processed by the April 2026 reporting, against roughly 45,000 calls connected to staff per month, in a record system that fed 31,390 SNAP application dispositions in December 2025 alone. No independent evaluation, inspector-general audit, or completed privacy impact assessment of the tool is on record.
empirical- Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
- Government Massachusetts Department of Transitional Assistance, DTA Performance Scorecard, December 2025 (2025) https://www.mass.gov/doc/performance-scorecard-december-2025-0/download
Oversight of the DTA call summarizer ran through the labor channel: SEIU Local 509, representing DTA call-center workers among roughly 9,000 state employees, reached an agreement with DTA — pre- or early-deployment; the record does not establish it preceded the December 2025 rollout — that made worker use voluntary and protected jobs, without addressing record provenance or client-side safeguards. The formal privacy apparatus sat empty: none of the nine AI use cases Massachusetts disclosed from its internal inventory of at least 40, the DTA summarizer included, reported a completed privacy impact assessment; the tool's interaction-data plan was listed as still to be defined while it was in production; and details on the other 31 use cases were withheld until the Supervisor of Records ordered them submitted for in camera review.
empirical- Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
- Government Massachusetts Executive Office of Technology Services and Security, Artificial Intelligence at the Commonwealth (Mass.gov) (2026) https://www.mass.gov/artificial-intelligence-at-the-commonwealth
The record store the AI summaries enter was itself contested while the pilot ran: in California v. USDA, No. 3:25-cv-06310 (N.D. Cal.), a 21-state-plus-DC coalition including Massachusetts obtained preliminary injunctions on October 15, 2025 and February 27, 2026 blocking USDA from cutting SNAP funding over states' refusal to hand over personal SNAP applicant and recipient data, the court holding the proposed data protocol would likely permit sharing beyond the entities allowed under 7 U.S.C. 2020(e)(8). Both orders are preliminary and the litigation is live; whether AI-generated call summaries held in BEACON fall within the demanded data is unresolved, and the Massachusetts AG's office declined to comment on that question.
empirical- Government Massachusetts Attorney General's Office, AG Campbell Secures Second Order Blocking Trump Administration From Cutting Off SNAP Funding Because of States' Refusal to Turn Over Personal Data of SNAP Applicants and Recipients (2026) https://www.mass.gov/news/ag-campbell-secures-second-order-blocking-trump-administration-from-cutting-off-snap-funding-because-of-states-refusal-to-turn-over-personal-data-of-snap-applicants-and-recipients
- Investigative JURIST, US federal court blocks SNAP funding cuts over states' refusal to share recipient data (2026) https://www.jurist.org/news/2026/02/us-federal-court-blocks-snap-funding-cuts-over-states-refusal-to-share-recipient-data/
- Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
Documented benefit-automation failures replicated determinations into downstream systems with no independent reconciliation against the source records — Michigan MiDAS actioned replicated flags and Robodebt reversed the onus onto recipients.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
- Government Royal Commission into the Robodebt Scheme, Report (2023) https://robodebt.royalcommission.gov.au/publications/report
- Investigative Law Society Journal, Crude, cruel and unlawful: Robodebt findings https://lsj.com.au/articles/crude-cruel-and-unlawful-robodebt-royal-commission-findings/
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Understand the system — Understand the system
- Check copied records — Reconcile copied records
- Mark AI-written records — Provenance labeling
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Review the riskiest first — Risk-tiered oversight
- Gate vendor updates — Vendor quality gate
- Review on schedule — Oversight cadence & retrospectives
Documented case histories
- Massachusetts DTA call summaries
- Magic Notes (Beam)
- Minute / Local Transcribe
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check