Skip to content

PAN Lab example

Nava assistive benefits chatbot

Done carefully: a verify-before-use copilot

For once, a system that starts safe. This copilot answers caseworker questions only from a vetted document set and hands back direct quotes to check — and it never writes to the record. Modeled on Nava's chatbot. The usual contamination barely exists here. The question flips: what keeps it that way when the pressure rises?

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Nava-class verify-before-use copilot network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the retrieval-grounded, verify-before-use assistive-copilot pattern documented in the Nava case file — not a reconstruction of the actual product.

  • assumed

    Peer pathways are authored on both signs: collective verification habits and citation checking are drawn as inhibiting peer-check pathways, and the dormant agent-to-agent coupling stays on the map so pressure can be seen opening it.

  • baseline

    The document set is curated and human-maintained, and the copilot does not write back to it, so record contamination is low by design at baseline.

  • assumed

    Contamination enters only if a caseworker adopts an answer without checking its cited source — the deference pathway this scenario watches.

What this example does not show

  • A safe starting baseline is a property of this model, not a safety promise for any real deployment.

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories