Skip to content

PAN Lab example

Imagine LA Benefit Navigator copilot

Best where you can check it least: a benefits-navigation copilot

This copilot answers a caseworker's benefits question from a curated set of approved government documents and hands back the exact quotes to check, all while the client is on the line. Modeled on the Imagine LA Benefit Navigator pilot. In the evaluation it helped most on the hardest questions and for the newest staff, which is the catch: the lift lands exactly where the person using it can check it least, and it fades as people stop engaging. The whole safety case is one habit — read the citation before you relay the answer — and habits decay.

Stylized model of a documented deploymentHousing & homelessness services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Benefit-Navigator-class assistive copilot network: 5 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the retrieval-grounded, verify-before-use assistive-copilot pattern documented in the Imagine LA Benefit Navigator case file — not a reconstruction of the actual chatbot. It is a deliberate sibling of the Nava-class verify-before-use copilot already in the library, and distinct from it: that case models Nava's earlier exploratory benefits-navigation work, while this one models the Gates-funded, co-evaluated Imagine LA pilot whose lift concentrated on the least-experienced staff.

  • baseline

    The documented accuracy lift concentrates on the hardest client questions and the newest, least-experienced staff, so the copilot becomes most load-bearing for exactly the caseworkers with the least independent capacity to catch its errors; the deference channel is drawn live at baseline to carry that dependence.

  • baseline

    Safety here is a maintained behavioral state, not a structural given: adoption was partial (about 65%) and declined over time without sustained engagement, so the verify-before-use habit and the engagement it rides on are treated as things that decay, and the collective peer check is drawn one step below the Nava-class copilot's.

  • assumed

    The retrieval corpus is a curated, human-maintained set of approved government documents and the copilot does not write back into it, so record contamination is low by design at baseline; contamination enters only if a caseworker adopts an answer without checking its cited source.

  • assumed

    Peer pathways are authored on both signs: relaying and adopted practice spread across the six pilot sites and the citation mechanism gives each answer an independent check, while the dormant agent-to-agent coupling stays on the map because the funded next step (agents that navigate portals and complete applications under supervision) is exactly what would open it.

  • assumed

    The independent evaluation is drawn as a periodic oversight inflow, not a per-case second read, and it was co-authored by the tool builder with its academic partners rather than independently replicated; the reported roughly 40% accuracy gain is a randomized-trial contrast on hypothetical client questions, not a live-caseload eligibility audit.

  • assumed

    Served families are not in the dynamics; the harm mode on this positive pole is omission or a wrong answer relayed to a client, recorded outside a diagram like this one, never a punitive flag. No differential client harm is computed here, and the documented readability gap (answers at a tenth-to-twelfth-grade reading level against college-level manuals) is an accessibility caveat, not a subpopulation disparity.

What this example does not show

  • Served families, and the benefits they do or do not ultimately receive, are not modeled here; the Lab models institutional propagation only, and those outcomes are documented in the case file and measured outside any diagram like this one.
  • The roughly 40% accuracy improvement is a randomized-trial contrast on hypothetical client questions, co-authored by the tool builder with its academic partners rather than independently replicated; it is decision-support accuracy, not a proven downstream enrollment effect, and the reported time-savings were explicitly promising but inconclusive.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The same evaluation reported that the chatbot's accuracy gains were largest on the most difficult client questions and among the newest, least-experienced staff (a directional finding, not a quantified breakdown), that about 65% of caseworkers with access used it at an average of about 14 prompts each and a modest, low-positive satisfaction (a Net Promoter Score of 11), that usage tended to decline over time without sustained engagement, and that answers averaged a tenth-to-twelfth-grade reading level against college-level source manuals.

    empirical
    • Vendor Nava Public Benefit Corporation (Nava Labs), Evaluating a GenAI-powered assistive chatbot for caseworkers (2026) https://www.navapbc.com/case-studies/evaluating-ai-assistive-chatbot-caseworkers
    • Academic Chen, Esposito, Giannella, Guo, Gosciak, Koenecke, Helping the Helpers: Evaluating a GenAI-powered assistive chatbot for caseworkers (Georgetown University Better Government Lab, Cornell University, and Nava PBC, 2026) https://digitalgovernmenthub.org/library/helping-the-helpers-evaluating-a-genai-powered-assistive-chatbot-for-caseworkers/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Housing & homelessness services domain page.

Levers available here and the patterns behind them

Documented case histories