PAN Lab example
Nava assistive benefits chatbot
Done carefully: a verify-before-use copilot
For once, a system that starts safe. This copilot answers caseworker questions only from a vetted document set and hands back direct quotes to check — and it never writes to the record. Modeled on Nava's chatbot. The usual contamination barely exists here. The question flips: what keeps it that way when the pressure rises?
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Nava-class verify-before-use copilot network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the retrieval-grounded, verify-before-use assistive-copilot pattern documented in the Nava case file — not a reconstruction of the actual product.
- assumed
Peer pathways are authored on both signs: collective verification habits and citation checking are drawn as inhibiting peer-check pathways, and the dormant agent-to-agent coupling stays on the map so pressure can be seen opening it.
- baseline
The document set is curated and human-maintained, and the copilot does not write back to it, so record contamination is low by design at baseline.
- assumed
Contamination enters only if a caseworker adopts an answer without checking its cited source — the deference pathway this scenario watches.
What this example does not show
- A safe starting baseline is a property of this model, not a safety promise for any real deployment.
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Keep prompts neutral — Framing and mirroring reduction
- Keep skills sharp — Deskilling-arrest mandate
- Gate vendor updates — Vendor quality gate
- Purge all old records — Content-aware decontamination
- Store less data — Data minimization
- Peer sharing rules — Peer-edge governance
- Assign a challenger — Structured dissent
- Escalate checks — State-feedback vigilance
Documented case histories
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down