PAN Lab example
Amsterdam Smart Check
The governed exit: a fair welfare screener that shipped every safeguard but one
This one did almost everything right. A city built a deliberately fair welfare-fraud screener: an explainable model with sensitive attributes and postal codes left out, a bias audit, debiasing that approximately equalized wrongful-flag rates between Western and non-Western applicants on past data, a data-protection assessment, a human-rights assessment, an external review, a citizen panel, and public transparency registers. Advisory only, with three separate humans in the loop. Modeled on Amsterdam's Slimme Check welfare-screening pilot. Then it went live, and the bias came back wearing the opposite face: now it wrongly flagged the majority group, women, and parents, flagged more applications than the old paper process, and was no better than the caseworkers. The city stopped it. Everything on this board was built before deployment. The one thing that was not is the thing that would have caught this: someone watching the live distribution, on a schedule, with the power to halt.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Amsterdam-Slimme-Check-class governed-exit screening pipeline network: 7 components and 15 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the governed-exit pattern documented in the Amsterdam Slimme Check case file — a best-practice, fully-assured screening pilot that the city itself halted after live bias re-emerged — not a reconstruction of the actual tool.
- assumed
Peer pathways are authored on both signs: the three-stage pipeline propagates a wrongly-validated label forward (validated label to investigation to final decision), the separated final-decision role cross-checks each case at a low level, and one screener scoring every application makes a distributional blind spot systematic — which is why three separate humans reading individual cases could not see it.
- assumed
Every recommended pre-deployment assurance step is present and actuated (the data-protection and human-rights assessments, the external review, the academic consultation, the two register entries, and the citizen panel), and the debiasing approximately equalized wrongful-flag rates between Western and non-Western applicants on retrospective data; the two defining absences are drawn inactive: the citizen panel's stop advice that was overridden, and the standing live-distribution re-audit that never existed.
- assumed
The model reads municipal records and income data as its features and is drawn here as logging an advisory score to the record — an inferred pathway: the documented surface is the advisory label shown to the reviewers, and no source records a score write-back; the human validates, investigates, and decides, and only a human write enters the permanent record, so whether error reaches a decision depends on the per-case check holding — a check that is structurally blind to the population-level error geometry.
- assumed
Served applicants, and whether they receive social assistance, are not in the dynamics. The demographic disparity figures are the investigating journalists' computations on city-provided aggregate confusion matrices under a GDPR remote-analysis regime, not city publications, and inverted in sign between the retrospective and live data; no differential client harm is estimated here, and those outcomes are documented in the case file and measured outside any diagram like this one.
What this example does not show
- Served applicants, and whether they receive social assistance, are not modeled here; the Lab models institutional propagation only, and those outcomes are documented in the case file and measured outside any diagram like this one.
- The demographic disparity figures (non-Western applicants nearly twice as likely wrongly flagged and men about 14 percent more before reweighting; women about 22 percent more, plus Dutch nationals and applicants with children, in the live pilot) are the investigating journalists' computations on city-provided aggregate confusion matrices under a GDPR remote-analysis regime, not city publications, and the sign inverted between the retrospective and live datasets; they are decision-support context, not causal findings, and no differential client harm is computed here.
- The roughly 535,000-euro cost and the pre-pilot claim of a 20 percent accuracy edge over caseworkers are the city's own internal estimates, not audited figures, and the accuracy edge did not hold in live deployment; the halt (register end date September 2023; public termination announced November 2023) is reported as the city's own decision, and the internal disagreement over whether stopping was a governance success or the abandonment of a fixable tool is carried in the case file, not adjudicated here.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The City of Amsterdam spent roughly five years and an estimated EUR 535,000 building a deliberately fair, explainable welfare-fraud screening model with nearly every recommended pre-deployment safeguard in place - a bias audit, training-data reweighting that approximately equalized wrongful-flag rates on retrospective data, a data-protection assessment and a human-rights assessment, external and academic review, a citizen panel, and dual algorithm-register transparency - and discontinued it after a 2023 live pilot on nearly 1,600 applications. In the investigating journalists' analysis of aggregate data the city provided, the group disparities re-emerged inverted on the live pilot, now more likely to wrongly flag Dutch nationals, women, and applicants with children, with the tool flagging more applications than the analog process and no better than caseworkers at finding genuine cases. The Dutch national algorithm register records the deployment ending September 2023 and lists it out of use, and the responsible alderman announced the halt in November 2023.
empirical- Investigative Braun, Geiger, Amsterdam Fair Welfare AI (Inside Amsterdam's high-stakes experiment to create fair welfare AI) (MIT Technology Review, with Lighthouse Reports and Trouw, 2025) https://www.technologyreview.com/2025/06/11/1118233/amsterdam-fair-welfare-ai-discriminatory-algorithms-failure/
- Investigative Lighthouse Reports, Amsterdam's 'Smart Check' welfare-fraud model: fairness methodology (with Trouw and MIT Technology Review, supported by the Pulitzer Center) https://www.lighthousereports.com/methodology/amsterdam-fairness/
- Government Algoritmeregister (Dutch national algorithm register), Onderzoekswaardigheid: Slimme check levensonderhoud (Gemeente Amsterdam) (2023, last modified 2025) https://algoritmes.overheid.nl/nl/algoritme/gm0363/95794697/onderzoekswaardigheid-slimme-check-levensonderhoud
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Escalate checks — State-feedback vigilance
- Review on schedule — Oversight cadence & retrospectives
- Keep skills sharp — Deskilling-arrest mandate
- Assign a challenger — Structured dissent
- Peer sharing rules — Peer-edge governance
- Check with a second model — Cross-model verification
- Vet connections — Connection authorization
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Mark AI-written records — Provenance labeling
- Pause AI on alarms — Deployment circuit-breaker
Documented case histories
- Amsterdam Smart Check
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot