PAN Lab example
Forsakringskassan VAB fraud-selection profile (Sweden)
The audit the agency refused: a secret fraud-selection profile no one outside could see
A secret machine-learning profile scores every applicant for a child-care benefit and routes the highest-scored to fraud investigators, and the outcomes of those investigations become the labels that train the next version. A separate random sample is investigated each year, producing bias-free ground truth about who is actually making errors, but that ground truth is never turned on the model, and the model, its features, and its precision are never shown to anyone outside. Modeled on Sweden's Forsakringskassan risk profile. Watch where the control has to come from: not from a better score, but from forcing the audit the agency refused, and from reconciling the profile against the ground truth it already had.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Forsakringskassan-class risk-based fraud-selection profile network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the risk-based fraud-selection pattern documented in the Sweden Forsakringskassan case file, not a reconstruction of the actual model or its undisclosed features. The disparity figures are outcome computations under specific fairness definitions from an obtained dataset, not confirmed model internals, and the agency disputed the framing.
- assumed
Peer pathways are authored on both signs: one selection profile's systematic skew and shared suspicion heuristics reinforce, while the inhibiting side is a standing review of the targeting that sits latent at baseline (the peer-review step a review lever would open).
- baseline
The feedback loop is present at baseline: past control outcomes become training labels, so groups historically over-investigated are selected again. The audit inspectorate (ISF, 2018) and the external investigation both flagged this data-bias dynamic qualitatively; the Lab models the loop's shape, not a longitudinally quantified rate.
- baseline
The random-control ground-truth labels are a real store the agency held: an unbiased arm (about 1,047 of 6,129 selections in the documented year) whose outcomes the profile's targeting could have been reconciled against. That reconciliation sits latent at baseline (a record-against-record check that carries nothing), reflecting that ground truth existed internally but was never turned on the model — the check the reconciliation lever supplies.
- baseline
The defining feature is an absence: no audit could see the model. The model, its features, and its precision were withheld under a roughly three-year freedom-of-information refusal, so both checks — reconciling the model against the random-control ground truth, and a standing review of who it flags — start closed at baseline, and the external-audit cadence and deployment gate were never in force. These are the controls the case never had, and the ones the disclosure and audit levers supply.
- assumed
The reported harm is demographic: the profile over-selected women, people of foreign background, below-median earners, and people without a university degree, and wrongly flagged those groups at higher false-positive rates. This Lab models institutional workflow propagation, not demographics, and estimates no differential harm to served people; those disparities are documented in the case file and measured outside any diagram like this one.
What this example does not show
- The reported harm is demographic: the profile over-selected women, people of foreign background, below-median earners, and people without a university degree, and wrongly flagged those groups at higher false-positive rates. The Lab models institutional workflow propagation, not demographics, and estimates no differential harm to served people; those disparities are documented in the case file and measured outside any diagram like this one.
- The disparity figures are outcome computations under specific fairness definitions, drawn from a single obtained year of outcome data, not confirmed model internals. The agency disputed the framing and never released the model, its features, or its precision, and no court or regulator issued a discrimination or GDPR finding: the data-protection regulator closed its supervision for mootness after the system was withdrawn. The Lab uses the case's shape, not calibrated rates.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Analysing the Swedish Social Insurance Agency (Forsakringskassan) 2017 outcome data, Lighthouse Reports and Svenska Dagbladet reported on 27 November 2024 that the agency's in-house machine-learning risk profile for the temporary parental allowance (VAB) selected women (more than 1.5x), people of a foreign background (about 2.5x), below-median earners (2.97x), and people without a university degree (3.31x) for fraud investigation more often than comparison groups by demographic parity, and wrongly flagged those groups at higher false-positive rates (about 1.7x for women and 2.4x for people of a foreign background); in the agency's paired random-control sample, 20.2 percent of applications contained at least one day incorrectly paid, an unbiased base error rate. These are outcome computations under specific fairness definitions from a single obtained year of data, not confirmed model internals; the agency disputed the framing and did not release the model. The data-protection regulator IMY closed its GDPR supervision on 18 November 2025 for mootness after the agency withdrew the system, and no court or regulator issued a discrimination or GDPR penalty.
empirical- Investigative Lighthouse Reports, Sweden's Suspicion Machine (co-published with Svenska Dagbladet, 27 Nov 2024) https://www.lighthousereports.com/investigation/swedens-suspicion-machine/
- Investigative Lighthouse Reports, How we investigated Sweden's Suspicion Machine (methodology) https://www.lighthousereports.com/methodology/sweden-ai-methodology/
- Investigative Lighthouse Reports, suspicion_machines_sweden (data and analysis repository, GitHub) https://github.com/Lighthouse-Reports/suspicion_machines_sweden
- Government Integritetsskyddsmyndigheten (IMY), Avslutad tillsyn efter att Forsakringskassan tagit AI-system ur bruk (Supervision closed after Forsakringskassan took AI system out of use) (2025) [Swedish] https://www.imy.se/nyheter/avslutad-tillsyn-efter-att-forsakringskassan-tagit-ai-system-ur-bruk/
- Government Integritetsskyddsmyndigheten (IMY), Tillsyn: Forsakringskassan (supervision case page and decision, 18 Nov 2025) (2025) [Swedish] https://www.imy.se/tillsyner/forsakringskassan/
The Swedish Social Insurance Agency (Forsakringskassan) did not disclose the machine-learning risk profile it used to select temporary-parental-allowance recipients for fraud investigation: its algorithm class, features, and precision were never released, and the agency resisted freedom-of-information disclosure for roughly three years on fraud-prevention grounds. In 2018 the audit inspectorate ISF found the risk-based profiling substantially more accurate than alternative controls while warning that it raised legal-certainty and equal-treatment concerns, and cautioning that an accurate model can still be inequitable when two groups err equally but only one is followed up. Amnesty International reported that a former agency data protection officer warned in 2020 that the operation breached European data-protection rules. The system was decommissioned in 2025 during the regulator's supervision, before any court or regulator ruled on it.
empirical- Investigative Lighthouse Reports, Sweden's Suspicion Machine (co-published with Svenska Dagbladet, 27 Nov 2024) https://www.lighthousereports.com/investigation/swedens-suspicion-machine/
- Investigative Lighthouse Reports, How we investigated Sweden's Suspicion Machine (methodology) https://www.lighthousereports.com/methodology/sweden-ai-methodology/
- Government evaluation Inspektionen for socialforsakringen (ISF), Profilering som urvalsmetod for riktade kontroller (Profiling as a selection method for targeted controls) (2018) [Swedish] https://isf.se/publikationer/rapporter/2018/2018-03-26-profilering-som-urvalsmetod-for-riktade-kontroller
- Government evaluation Inspektionen for socialforsakringen (ISF), Riskbaserade urvalsprofiler och likabehandling (Risk-based selection profiles and equal treatment) (2018) [Swedish] https://isf.se/publikationer/rapporter/2018/2018-06-15-riskbaserade-urvalsprofiler-och-likabehandling
- Advocacy Amnesty International, Sweden: Authorities must discontinue discriminatory AI systems used by welfare agency (2024) https://www.amnesty.org/en/latest/news/2024/11/sweden-authorities-must-discontinue-discriminatory-ai-systems-used-by-welfare-agency/
Where this connects
Institutional pressures in this domain
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Public benefits & eligibility domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Require sign-off — Conformity assessment gate
- Mark AI-written records — Provenance labeling
- Assign a challenger — Structured dissent
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Understand the system — Understand the system
- Upgrade model — Improve the model
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
Documented case histories
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Michigan MiDAS
- Robodebt (Australia)
- Indiana / IBM eligibility modernization
- Rotterdam welfare-fraud risk model
- Arkansas ARChoices / ARIA
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- SyRI (Netherlands)
- CNAF benefit-fraud risk score (France)
- Udbetaling Danmark data-driven control (Denmark)
- BOSCO (Spain)
- Serbia Social Card (Socijalna karta)
- UK DWP Universal Credit Advances fraud model
- ID.me identity verification as an unemployment eligibility gate
- Medicaid unwinding: automated ex parte renewal at population scale
- INSS auto-analysis: when the productivity metric makes denial the fastest way out
- Samagra Vedika
- Workforce Australia Targeted Compliance Framework: automated payment sanctioning after Robodebt
- NYC MyCity business chatbot
- Nevada DETR generative-AI unemployment appeals
- Tennessee TennCare TEDS