PAN Lab example
The same AI running hands off: the agentic office
The AI here doesn't just draft casework — it acts on cases, and a stretched staff waves most of it through. Same model as the other two offices. In the sociotechnical simulation — a modeled office, not a real one — errors stick here about 75% of the time, versus 20% and 16% next door. Find the levers that change that.
Open this example in the PAN Lab to apply pressures and levers and watch what the system does.
What this models
This example runs on the Agentic low-oversight office network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
- baseline
How often errors stick in this office (≈75%) comes from the sociotechnical simulation, not from measuring a real office.
- assumed
The model's raw error output is held identical across all three office cultures; only oversight differs.
- assumed
Verification behavior is treated as homogeneous within the office.
- assumed
Peer pathways are authored on both signs: agent-to-agent chaining and coworker contagion run at baseline, while the inhibiting peer checks (second opinions, cross-model verification) start closed — this culture's defining absence.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In the sociotechnical simulation, the same AI in three modeled office cultures - stylized, not real workplaces - let errors stick at very different rates: roughly 75% under low-oversight autonomy, 20% under human supervision, and 16% under high-governance professional controls.
scenarioillustrative PAN-run resultNo published source is attached to this claim yet.
Where this connects
Institutional pressures in this domain
- Caseload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Verify output — Put a verifier on the agent
- Upgrade model — Improve the model
- Gate record entries — Human-in-the-loop write gating
- Mark AI-written records — Provenance labeling
- Escalate checks — State-feedback vigilance
- Pause AI on alarms — Deployment circuit-breaker
- Vet connections — Connection authorization
- Peer sharing rules — Peer-edge governance
- Purge all old records — Content-aware decontamination
- Keep prompts neutral — Framing and mirroring reduction
- Store less data — Data minimization
- Assign a challenger — Structured dissent
- Check with a second model — Cross-model verification
Documented case histories
- Magic Notes (Beam)
- Minute / Local Transcribe
- Justice Transcribe
- GDS Microsoft 365 Copilot cross-government experiment
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check