Skip to content

PAN Lab example

Rotterdam welfare-fraud risk model

The suspicion machine: a welfare-fraud risk model

A vendor model ranks welfare recipients by fraud risk, and investigators work the list top-down. Modeled on Rotterdam's system. Watch the retraining loop: the model learns from the outcomes its own rankings produced, so a skew in who gets investigated rides straight back into who it suspects next.

Stylized model of a documented deploymentPublic benefits & eligibility

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Rotterdam-class fraud risk-scoring model network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the machine-learning welfare-fraud risk-tool pattern documented in the Rotterdam case file — not a reconstruction of the actual model or its inputs.

  • assumed

    Peer pathways are authored on both signs: shared investigation heuristics and a single model's systematic skew reinforce, while the internal peer-challenge pathway starts closed — the documented scrutiny came from outside auditors.

  • baseline

    The retraining loop — investigators' labeled outcomes and historical dossiers feeding the next model — is present at baseline, reflecting the audit's account of a model trained on administrative and outcome data.

  • assumed

    The external-data-sources node carries a real inflow into the model — the roughly 315 municipal and demographic inputs the Suspicion Machines audit documented the model being built on and scoring from — so it enters the dynamics, not just the picture. What those inputs encode as differential harm stays external (below).

  • assumed

    The ranked risk list is drawn as a mediating artifact on the model → investigators pathway — reflecting the case file's account of a system whose output was a prioritized list of people to investigate. It carries no flow of its own and does not affect the dynamics.

  • assumed

    Documented demographic disparities in who this kind of model flags are recorded externally in the case file. This Lab models institutional propagation, not demographics, and estimates no differential harm to served people.

What this example does not show

  • The documented harm is a demographic disparity in who was flagged. The Lab models institutional workflow propagation, not demographics, and estimates no differential harm to served people; that disparity is documented in the case file and measured outside any diagram like this one.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.

    empirical
    • Government Rekenkamer Rotterdam, Gekleurde technologie: onderzoek naar het gebruik van algoritmes door de gemeente Rotterdam (2021) https://www.rekenkamers.nl/rapport/gekleurde-technologie/
    • Investigative Lighthouse Reports, Suspicion Machines (2023) https://www.lighthousereports.com/investigation/suspicion-machines/
    • Investigative WIRED / Lighthouse Reports, Inside the suspicion machine (2023) https://www.wired.com/story/welfare-state-algorithms/
    • Investigative Follow the Money, How a fraud algorithm learned to suspect vulnerable groups https://www.ftm.eu/articles/algorithm-rotterdam-dissected
    • Advocacy Racism and Technology Center, Rotterdam welfare fraud algorithm was biased https://racismandtechnology.center/2023/03/17/racist-technology-in-action-rotterdams-welfare-fraud-prediction-algorithm-was-biased/
  • Documented risk-scoring deployments computed scores from multi-agency administrative records originally collected for other purposes, which is the data-protection critique recorded in independent reviews of these systems.

    empirical
    • Government evaluation Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf
    • Investigative Eubanks, Automating Inequality (2018); The Nation, Want to Cut Welfare? There's an App for That https://www.thenation.com/article/archive/want-cut-welfare-theres-app/
    • Investigative Lighthouse Reports, Suspicion Machines (2023) https://www.lighthousereports.com/investigation/suspicion-machines/

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Public benefits & eligibility domain page.

Levers available here and the patterns behind them

Documented case histories