PAN Lab example
The same AI with a human checking: the supervised office
Same AI — but a human checks every output before it reaches the record. In the sociotechnical simulation — a modeled office, not a real one — that holds errors to about 20%, a quarter of the hands-off rate. The failure here is slow: the review keeps happening while the judgment behind it quietly thins.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Human-supervised office network: 3 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- baseline
How often errors stick in this office (≈20%) comes from the sociotechnical simulation, not from measuring a real office.
- assumed
Peer pathways are authored on both signs: answers spread collegially and reviewers compare notes; the shared assistant's errors are treated as correlated across the team (a monoculture assumption, not a measurement).
- assumed
The model's raw error output is held identical across all three office cultures; only oversight differs.
- assumed
Review quality is assumed steady over time; the deskilling drift explored elsewhere is switched off at baseline.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In the sociotechnical simulation, the same AI in three modeled office cultures - stylized, not real workplaces - let errors stick at very different rates: roughly 75% under low-oversight autonomy, 20% under human supervision, and 16% under high-governance professional controls.
scenarioillustrative PAN-run resultNo published source is attached to this claim yet.
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Review the riskiest first — Risk-tiered oversight
- Keep prompts neutral — Framing and mirroring reduction
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Understand the system — Understand the system
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Store less data — Data minimization
- Assign a challenger — Structured dissent
Documented case histories
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check