PAN Lab example
Justice Transcribe
The note that scores you: a copilot at the head of a risk pipeline
A copilot transcribes probation supervision sessions and drafts the summary that becomes the case record - and it decides nothing. But the same record is read, at more than a thousand assessments a day, by a separate tool that scores reoffending risk, and an officer acts on that score. Modeled on the Ministry of Justice's Justice Transcribe. The trap is the seam: a note written to save ten minutes becomes an unchecked input to a risk score no one reconciled it against. Watch the lineage, and the fact that the tool reached every officer before anyone published how often it is right.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Justice-Transcribe-class multi-hop lineage copilot network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the multi-hop lineage pattern documented in the Justice Transcribe case file - not a reconstruction of the actual tool. The coupling between the copilot's output and the downstream risk profiler is modeled structurally, via a shared case-record ecosystem, not as a verified named data pipeline; no source documents an explicit data flow from the copilot's output fields into the risk tool's inputs.
- baseline
The defining dynamic is a two-model lineage: one copilot writes the statutory supervision record (writing to the record behind a review-and-accept step), and a second, higher-stakes algorithmic component - the ministry's reoffending-risk profiler, reporting places above 1,300 assessments a day - reads that same record as its scoring substrate (the record-to-model lineage), then the officer acts on the score. A copilot-written line is therefore not a local convenience but an unmeasured input to a high-stakes risk score. The reconciliation across that seam is latent - the case's signature absence.
- baseline
The second defining feature is that scaling outran evaluation: the tool went from pilot to every probation officer in England and Wales in about a year, with meeting volume roughly quadrupling between the two published transparency windows, while no transcription-accuracy, error-rate, or officer-correction evaluation was published and no independent evaluation exists. That is the latent accuracy-check pathway (governance to staff) - departmental governance publishes usage and satisfaction, not accuracy - opened by a scheduled challenge duty or a review cadence.
- assumed
The peer pathway is authored open: consistent summaries that follow individuals across practitioners are marketed as a benefit, but one shared summarisation layer homogenizes the record's phrasing across the workforce and summary and prompt habits spread officer to officer. All three check pathways start closed (the cross-lineage reconciliation, the standing accuracy audit, and the reconciliation before a recall or court report); the only scrutiny that exists is parliamentary, academic and civil-society, not a tool-specific accuracy audit.
- assumed
The downstream risk model's documented lower predictive validity for some ethnic groups is a property of that risk model, not of this copilot. The ministry's own validation found lower predictive validity for all Black, Asian and Minority Ethnic groups for non-violent reoffending, and for Black and Mixed ethnicity offenders for violent reoffending. Those served people are not on this diagram and never in these dynamics; that finding, and any pattern in who it affects, are recorded in the case file and measured outside any diagram like this one.
- assumed
Every time-savings figure cited for this tool (a 50 percent note-taking reduction, up to 240,000 days a year, a modelled 450,000 hours a year restated as about 18,750 calendar days, an illustrative ten-minutes-per-meeting assumption) is self-reported by users or asserted by government, not independently measured, and no transcription-accuracy evaluation is published; this Lab models institutional propagation only and estimates none of them.
What this example does not show
- The downstream reoffending-risk model has documented lower predictive validity for some ethnic groups, outcome-specific in the ministry's own validation. That is a property of the risk model, not of this copilot. The Lab models institutional propagation only, no demographics and no differential harm to served people; that finding is recorded in the case file and measured outside any diagram like this one.
- The coupling from the copilot's output to the downstream risk model is modeled structurally, via a shared case-record ecosystem, not as a verified named data pipeline; no source documents an explicit data flow from the copilot's fields into the risk tool's inputs. Every time-savings figure cited for this tool is self-reported or asserted by government, not independently measured, and no transcription-accuracy evaluation is published. This shape models how the tool moves through probation workflow and estimates none of those quantities.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The Ministry of Justice built an in-house AI transcription and summarisation copilot, Justice Transcribe, for probation staff in England and Wales, scaling it from a pilot to more than 1,000 officers in October 2025 and to every probation officer by June 2026, with official transparency data recording more than 800,000 supervision meetings summarised between 7 October 2025 and 2 June 2026; the reported time-savings are the ministry's own and rest on an operating assumption the department itself labels illustrative, and no transcription-accuracy rate, officer correction rate, or independent evaluation of the tool has been published.
empirical- Government Justice AI Unit, Ministry of Justice, Justice Transcribe in Probation (2026) https://ai.justice.gov.uk/our-work/justice-transcribe
- Government Ministry of Justice, AI Action Plan for Justice (GOV.UK, 2025) https://www.gov.uk/government/publications/ai-action-plan-for-justice/ai-action-plan-for-justice
- Government Ministry of Justice and HM Prison and Probation Service, Justice Transcribe data 7 October 2025 to 2 June 2026 (transparency data, GOV.UK, 2026) https://assets.publishing.service.gov.uk/media/6a1eafe265bc5f798327f61f/Justice-transcribe-report-2-june-2026.pdf
- Government Ministry of Justice and DSIT, OpenAI to expand into UK data hosting after major growth deal (GOV.UK press release, 2025) https://www.gov.uk/government/news/openai-to-expand-into-uk-data-hosting-after-major-growth-deal
- Government Ministry of Justice, AI tech ambition to deliver smarter justice for victims (GOV.UK press release, 2026) https://www.gov.uk/government/news/ai-tech-ambition-to-deliver-smarter-justice-for-victims
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
- Vet connections — Connection authorization
- Escalate checks — State-feedback vigilance
- Keep prompts neutral — Framing and mirroring reduction
- Gate record entries — Human-in-the-loop write gating
- Mark AI-written records — Provenance labeling
- Understand the system — Understand the system
- Store less data — Data minimization
- Peer sharing rules — Peer-edge governance
- Upgrade model — Improve the model
Documented case histories
- Justice Transcribe
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check