PAN Lab example
DWP Whitemail Insights and Vulnerability Scanner
The letter no one reads twice: an upstream vulnerability scanner
This scanner reads every one of roughly 25,000 paper letters a day sent to the welfare agency and decides which claimants surface on the potentially-vulnerable shortlist. Modeled on the Department for Work and Pensions (DWP) Whitemail Insights and Vulnerability Scanner. The transparency record names precision, recall and F1-score as its metrics and discloses no values; no independent evaluation exists. Nothing decides anything here: a flagged case gets a trained caseworker who can pull the original letter to check it, but a letter the scanner does not flag simply stays in the ordinary queue, read once and never again for vulnerability. The claimants were never told the tool reads their post - the impact assessment said they do not need to know - so a missed flag files no complaint. Watch the asymmetry, where the positives get a second read and the negatives get none, and the one number that would matter most: how often anyone actually pulls the letter.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Whitemail-scanner-class upstream correspondence triage network: 6 components and 15 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the upstream-correspondence-triage pattern documented in the Department for Work and Pensions (DWP) Whitemail Insights and Vulnerability Scanner case file - not a reconstruction of the actual tool. It is deliberately distinct from the library's other copilots: the Magic-Notes-class documentation writer sits on the error-to-record pathway with one drifting review gate; the Nava-class and Benefit-Navigator-class tools are verify-before-use cited-answer copilots; the Cross-government-copilot-class is a horizontal productivity layer whose lesson is measurement-is-not-control; and the Minute-class scribe's split is governance-versus-accuracy across a cohort. This shape is an upstream sensor that acts only on what it flags.
- baseline
The defining dynamic is an asymmetry around an invisible error. Human oversight is real but one-sided: a flagged case is assessed one at a time and the caseworker can retrieve the source letter image to verify it, but a document the scanner does not flag stays in the ordinary queue with no second read for vulnerability, so discretion is high on the positives the tool surfaces and effectively absent on the negatives it drops. The two latent check pathways - a standing second read of the residual, and an independent accuracy read of the scanner's outputs - are the absences levers can open.
- baseline
The invisible-error framing is grounded but partly analytical: it is documented that letter writers are not told the tool reads their correspondence (the data protection impact assessment reported by the Guardian stated they do not need to know), and that prioritising some cases inevitably deprioritises others (Turn2us, in the Guardian); the inference that a missed flag therefore generates no complaint and surfaces only as downstream harm attributed elsewhere is the Lab's reading of that documented design, not a claim any source states verbatim, and no such harm has been adjudicated.
- assumed
The source-image verification affordance is the system's true safety parameter and it is unmeasured: the unique document identifier lets a caseworker retrieve the original letter to verify a flag, but the rate at which anyone actually does so, the override or disagreement rate, and any false-negative estimate for the vulnerability channel are all unpublished. The base error rate is likewise an assumed modelling choice: the transparency record names precision, recall and F1-score as its metrics but discloses no values, and no independent evaluation exists.
- assumed
There is no egress pathway on this diagram, and that is a documented feature rather than an omission: the system is hosted with end-to-end encryption and no internet connectivity, personal data is automatically redacted after scanning, and there is no commercial data-processing agreement to an outside processor (the vendor built the tool and the Department for Work and Pensions (DWP) retains the intellectual property). The load-bearing risk here is the invisible miss, not a data leak, which is what leaves the accuracy and correction questions as the ones that matter.
- assumed
Peer pathways are authored on both signs: working-practice and threshold habits spread desk to desk, one central scanner reading every letter homogenises a systematic blind spot across all of them at once (a monoculture assumption in the Lab's qualitative vocabulary, not a measurement), and the daily manual review is a real, active team check - while the two accuracy-and-correction checks start closed, because the only external scrutiny is independent analyst critique and parliamentary correspondence, not a commissioned evaluation.
- assumed
The potentially vulnerable claimants a missed flag would leave in the ordinary queue are not on this diagram and never in these dynamics. Whatever a miss means for the person who wrote the letter, and any pattern in who is affected (the Equality Analysis exists but is unpublished), is documented in the case file and measured outside any diagram like this one; this Lab models institutional propagation only, and every performance and time-saving figure cited for the tool is an operator or government claim, not a measured quantity.
What this example does not show
- The potentially vulnerable claimants a missed flag would leave in the ordinary queue are not modeled here; the Lab models institutional propagation only, with no demographics and no differential harm to served people. An Equality Analysis exists but is unpublished, so any pattern in who a miss affects is documented in the case file and measured outside any diagram like this one.
- No accuracy figures are published for this tool: the transparency record names precision, recall and F1-score as its metrics but discloses no values, no independent evaluation exists, and the source-image verification usage rate, the override or disagreement rate, and any false-negative estimate for the vulnerability channel are all unpublished. The invisible-error framing is the Lab's analytical reading of the documented non-notification and shortlist design; media contestation of the tool is reported concern, not adjudicated harm.
- Every throughput and time-saving figure is an operator or government claim, not an independently measured quantity: daily volume is reported as 22,000 (end-2023 and March 2024) rising to circa 25,000 (November 2025), and the four-to-six-weeks-to-75-per-cent-same-day improvement comes from a Department for Work and Pensions (DWP) profile with no independent verification. Supplier attribution is contested - the transparency record names Accenture (UK) Limited, while earlier freedom-of-information (FOI)-based research suspected a different supplier - and the base error rate in this shape is an assumed modelling choice, not a calibrated value.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
According to its Algorithmic Transparency Recording Standard record, published on November 27, 2025, the UK Department for Work and Pensions runs a Whitemail Insights and Vulnerability Scanner that reads roughly 25,000 scanned documents a day (reported as around 22,000 a day at end-2023 and in a March 2024 operator interview). Each document is passed first through the Vulnerability Scanner, a pre-trained open-source transformer doing zero-shot classification, which flags potentially vulnerable customers against eight prescribed themes including suicide and self-harm, domestic violence and abuse, and financial hardship; only documents not flagged as indicating vulnerability are relayed to Whitemail Insights for routing across nine themes. The output to trained staff is an anonymised daily report of flagged customers, and DWP states the tool does not make or influence benefit entitlement decisions. The record names precision, recall, and F1-score as its evaluation metrics but discloses no values, and no independent accuracy evaluation has been published.
empirical- Government Department for Work and Pensions, Algorithmic Transparency Record: Whitemail Insights and Vulnerability Scanner (GOV.UK, 2025) https://www.gov.uk/algorithmic-transparency-records/whitemail-insights-and-vulnerability-scanner
- Trade press Trendall, DWP taps AI to scan 25,000 letters a day and identify vulnerable citizens (PublicTechnology, 2025) https://www.publictechnology.net/2025/12/08/society-and-welfare/dwp-taps-ai-to-scan-25000-letters-a-day-and-identify-vulnerable-citizens/
- Government UK Parliament Work and Pensions Committee, DWP use of artificial intelligence: correspondence (2023) https://committees.parliament.uk/publications/42458/documents/211057/default/
- Trade press Corbridge (interview), How DWP is getting AI to work (Computing, 2024) https://www.computing.co.uk/interview/4188076/dwp-getting-ai
Guardian FOI reporting in January 2025 recorded that benefit claimants are not told the AI reads their correspondence: the internal data protection impact assessment stated that letter writers do not need to know about their involvement in the initiative, and the tool had been piloted since at least 2023 without appearing on the central government AI transparency register despite a ministerial mandate. The correspondence it processes can include national insurance numbers, health information, bank details, and children's details. Turn2us policy manager Meagan Levin voiced serious concerns, noting that prioritising some cases inevitably deprioritises others, so it is vital to understand how these decisions are made and ensure they are fair. The further reading that a missed flag on the unflagged residual therefore has no complaint channel and surfaces only as downstream harm is an analytical inference from the documented non-notification and shortlist design, not an adjudicated harm.
empirical- Investigative Booth, Serious concerns about DWP use of AI to read correspondence from benefit claimants (The Guardian via inkl, 2025) https://www.inkl.com/news/serious-concerns-about-dwp-s-use-of-ai-to-read-correspondence-from-benefit-claimants
- Trade press Toth, AI use for welfare system in doubt as scale of DWP setbacks revealed (The Independent via Yahoo News, 2025) https://www.yahoo.com/news/ai-welfare-system-doubt-scale-170440585.html
- Advocacy Dent, Digital Welfare State edition 006 (ABD Consultancy, 2025) https://www.abdconsultancy.co.uk/blog/digitalwelfarestateedition006
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Check with a second model — Cross-model verification
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Vet connections — Connection authorization
- Peer sharing rules — Peer-edge governance
- Mark AI-written records — Provenance labeling
- Upgrade model — Improve the model
Documented case histories
- DWP Whitemail Insights and Vulnerability Scanner
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check