PAN Lab example
UK Home Office asylum AI copilots
Compressed before judgment: the summary no one is told to check
Two copilots sit in asylum casework: one compresses the interview transcript — the claimant's own account — into a summary the decision-maker reads, and one summarises the country policy at decision time. Modeled on the Home Office asylum AI trials. In the pilot the summariser saved a measured 23 minutes a case, 9% of its summaries were removed as inaccurate before any caseworker saw them, and its summaries carried no source references back to the transcript. Nothing here decides a claim; the tools are advisory and the transcript stays available. But the 23 minutes saved are the minutes not spent re-reading, no one is required to check a summary, and the pilot filter that caught the 9% has no documented production version. Watch the compression edge where evidence is squeezed before judgment, and watch who is left to catch a wrong summary — because the one person who would know, the claimant, is never told AI touched their case.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Asylum-summarisation-class compression-edge copilot network: 7 components and 16 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the compression-edge summarisation pattern documented in the Home Office asylum AI copilots case file — not a reconstruction of the actual Asylum Case Summarisation or Asylum Policy Search tools or of any deployment. It is deliberately distinct from the library's other copilots: the Magic-Notes-class documentation tool rests its whole safety case on one human review gate on the error-to-record pathway; the Nava-class and Benefit-Navigator-class tools are verify-before-use, with a caseworker reading a cited answer; and the Cross-government-copilot-class is a horizontal layer whose lesson is measurement-is-not-control. Here two vertical compression edges sit between the evidence record and the human judgment node in a decision with an asymmetric, near-irreversible downside.
- baseline
The two compression edges encode the case's substrate: the summariser compresses the interview transcript (the primary evidence of the claimant's account) and the policy-search tool compresses the Country Policy Information corpus. Compression-of-evidence-before-judgment is the substrate, not an asserted defect — the tools are advisory and the transcript remains available.
- baseline
The distinctive dynamic is that the throughput gain is the exact adoption incentive that erodes per-case verification: the pilot measured a 23-minute-per-case time saving (a 32% reduction) for the summariser, and those minutes exist only insofar as the decision-maker does not redo the reading the summary replaced. Verification is unpriced — there is no requirement to check summaries against the transcript, and the pilot summaries lacked source references, raising the cost of any check — so the store-to-operator read-back is held low and the deference channel runs at baseline into both edges.
- baseline
The defining absence is drawn twice, both latent, and named once more in prose. On the diagram: no independent check verifies a summary against the source transcript before it enters the record (the model-side check starts closed), and the pilot's pre-use accuracy filter — the layer that removed 9% of summaries as inaccurate before caseworkers saw them, and the source of that measurement — has no documented standing production equivalent (the oversight node's check runs dormant). The deepest absence cannot be drawn: the one party with first-hand knowledge of whether the summary of their own account is right, the claimant, is severed from the loop by non-disclosure, so the only remaining detector is the operator whose verification the throughput incentive is designed to reduce.
- assumed
The 9% is a pilot pre-use filter rate — summaries removed as inaccurate or incomplete before any caseworker saw them — not a rate of errors reaching decisions, and its production persistence is undocumented. The model's raw error rate here is a modeling choice, not a measured per-interaction rate; the Home Office's own Calibre quality-assurance reviews found no statistically significant difference in decision quality on small pilot samples, a self-evaluation treated as an operator claim of quality neutrality, not an independent validation.
- assumed
Peer pathways are authored on both signs: habits of relying on the summary spread desk to desk, colleagues second-read each other's summaries against the transcript when time allows, and one shared summarisation tool compresses every asylum interview so a systematic mis-summarisation mode recurs across every case at once — a monoculture assumption in the Lab's qualitative vocabulary, not a measurement.
- assumed
The asylum claimants these decisions concern are not in the dynamics. The real-world stakes of a wrong decision — a wrongful refusal of protection, potentially refoulement — are documented in the case file and measured outside any diagram like this one; this Lab models institutional propagation only. Error distribution across nationalities and languages is undisclosed, so no differential harm is estimated. The commissioned legal opinion that the deployment is likely unlawful is a contested legal position, not a court ruling.
What this example does not show
- The asylum claimants these decisions concern are not modeled here, and neither is the real-world stakes of a wrong decision — a wrongful refusal of protection, potentially refoulement. The Lab models institutional propagation only; those outcomes are documented in the case file and measured outside any diagram like this one. Error distribution across nationalities and languages is undisclosed, so no differential harm is estimated.
- The 9% figure is a pilot pre-use filter rate — summaries deemed inaccurate or incomplete and removed before caseworkers saw them during the pilot — not a rate of errors reaching decisions, and no equivalent production filter, nor any post-deployment error, override, or monitoring data, is published. The Home Office's own Calibre quality-assurance reviews found no statistically significant difference in decision quality on small pilot samples: a self-evaluation treated here as an operator claim of quality neutrality, not an independent validation.
- The published evaluation recommended addressing the identified limitations before a full rollout, continuous monitoring in early rollout, and a larger-scale evaluation after deployment; the expansion was announced alongside the evaluation, not in later defiance of it. As of mid-2026 no post-deployment evaluation or monitoring data, and no published Data Protection Impact Assessment, Equality Impact Assessment, or Algorithmic Transparency Recording Standard entry, is in the record. The commissioned KC legal opinion that the deployment is likely unlawful on procedural-fairness and data-protection grounds is a contested legal position, not a court ruling; no litigation outcome exists as of mid-2026.
- Underlying model and vendor details are reported by civil society, not confirmed in the official evaluation (which names only a Large Language Model); no AI model identifier is used here. The per-case time figures (about 23 minutes and a 32% reduction for the summariser, about 37 minutes for the policy-search tool) are the pilot evaluation's own measured figures; the raw pilot case counts reflect unequal group sizes and are not evidence of per-worker throughput multiplication, so only the per-case time figures are used.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In the UK Home Office's own pilot of an AI tool that summarises asylum interview transcripts for decision-makers, 9% of the generated summaries were deemed inaccurate or incomplete and removed by a pre-use filter before any caseworker saw them, and 23% of users reported not being fully confident in the rest; the summaries carried no source references back to the transcript. The official evaluation, published April 29, 2025, measured a 23-minute-per-case time saving (a 32% reduction) for the summariser and about 37 minutes for a companion policy-search tool, and Home Office Calibre quality-assurance reviews found no statistically significant difference in decision quality on small pilot samples. The evaluation recommended addressing the identified limitations before a full rollout, continuous monitoring in early rollout, and a larger-scale evaluation after deployment; the Home Office announced expansion the same day. By January 2026 the policy-search tool had been rolled out to all asylum decision-makers, and per trade-press reporting the summarisation tool entered national rollout in April 2026.
empirical- Government evaluation UK Home Office (GOV.UK), Evaluation of AI trials in the asylum decision making process (2025) https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process
- Advocacy Open Rights Group, Saving time, risking lives: government uses AI tools to inform asylum decisions (2026) https://www.openrightsgroup.org/blog/saving-time-risking-lives-government-uses-ai-tools-to-inform-asylum-decisions/
- Trade press Government Transformation, Second AI tool for asylum caseworkers to be rolled out this month (2026) https://www.government-transformation.com/data/second-ai-tool-for-asylum-caseworkers-to-be-rolled-out-this-month
The Home Office's asylum interview-summarisation tool inserts a compression step whose measured value is a 23-minute-per-case time saving that exists only insofar as the decision-maker does not redo the reading the summary replaced: caseworkers are not required to verify summaries against transcripts, and the pilot summaries carried no source references that would make checking cheap. The correction loop is also severed from the other side. In a May 2026 written parliamentary answer, minister Alex Norris confirmed that asylum claimants are not told about the AI tools used in their cases, so the one party with first-hand knowledge of their own account cannot surface a summary error; this postdates Article 22C of UK GDPR (in force February 5, 2026). As of mid-2026 the rollout had proceeded without a published post-deployment evaluation or continuous-monitoring data and, per Open Rights Group, without a published Data Protection Impact Assessment, Equality Impact Assessment, or Algorithmic Transparency Recording Standard entry, with prompts withheld under a Freedom of Information refusal. A March 16, 2026 commissioned legal opinion argues the use is likely unlawful on procedural-fairness and data-protection grounds; that is a contested legal position, not a court ruling.
empirical- Government evaluation UK Home Office (GOV.UK), Evaluation of AI trials in the asylum decision making process (2025) https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process
- Trade press ResultSense, Home Office withholds AI details from asylum claimants (2026) https://www.resultsense.com/news/2026-05-07-home-office-asylum-ai-transparency/
- Advocacy Open Rights Group, Automating the hostile environment: AI in the asylum decision making process (2026) https://www.openrightsgroup.org/publications/automating-the-hostile-environment-ai-in-the-asylum-decision-making-process/
- Advocacy Open Rights Group, Home Office use of AI in asylum cases likely to be unlawful, legal opinion finds (2026) https://www.openrightsgroup.org/press-releases/home-office-use-of-ai-in-asylum-cases-likely-to-be-unlawful-legal-opinion-finds/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Check with a second model — Cross-model verification
- Review on schedule — Oversight cadence & retrospectives
- Escalate checks — State-feedback vigilance
- Gate record entries — Human-in-the-loop write gating
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Vet connections — Connection authorization
- Store less data — Data minimization
- Peer sharing rules — Peer-edge governance
- Upgrade model — Improve the model
Documented case histories
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check