Skip to content

PAN Lab example

CNAF benefit-fraud risk score (France)

The score that suspects the vulnerable: a benefit-fraud risk model

A national model scores every benefit-receiving household for fraud suspicion each month, and controllers work the highest scores first — up to the most invasive on-site checks. Modeled on France's CNAF/CAF datamining score. Watch two things at once: the variables that raise the score are the markers of the population the safety net exists to serve (low income, disability allowance, single parenthood, a recent separation or move), and the model is retrained on the outcomes of its own controls — so a skew in who gets checked rides straight back into who it suspects next.

Stylized model of a documented deploymentPublic benefits & eligibility

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the CNAF-class benefit-fraud risk-scoring model network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 6 assumed · 4 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • baseline

    The documented oversight ensemble is drawn as a wired-in reviewer rather than left in comments: the freedom-of-information body ordered the source code released, the ombudsperson filed observations to the administrative court in October 2025, and a coalition that grew to 25 organisations brought an annulment case in October 2024. The inbound pathway is the model itself (successive versions obtained under freedom-of-information law; the 2018 version published January 2026), drawn faintly because disclosure came through multi-year proceedings rather than routine practice. The outbound review of who controllers select is drawn empty: the scrutiny reached the model - a redesigned version entered production in January 2026 with the disability-allowance-while-working, nationality, housing-type and behavioural variables removed - but never became a scheduled review of selection, and no court had ruled.

  • baseline

    The routing edge is drawn at full strength and the two training-side arms are re-derived from the verified methodology. Nearly seven in ten investigations start from a household the score selected, which makes the routing channel dominant here. The training labels come from the annual random-audit survey, which the methodology reports as covering 98% of fraudulent payments, so that input carries full weight; controller determinations shaping the fitted outcome variable are drawn faint, because labels flowing back from targeted controls are an interpretive reading of the record rather than a documented mechanism.

  • assumed

    The documented methodology names a second input, so it is drawn: the annual random-audit survey that supplies the model's ground truth and its baseline yield. It is the one part of this pipeline that looks at households selected by chance rather than by the score's own history - which is what makes it ground truth at all, and also what makes it structurally different from the household data the score reads.

  • assumed

    This models the national benefit-fraud risk-scoring (datamining) pattern documented in the France CNAF/CAF case file — not a reconstruction of the actual model, its variables, or its coefficients.

  • assumed

    Peer pathways are authored on both signs: one national score's systematic skew and shared controller targeting heuristics reinforce, while a real human check survives — controllers make the final fraud-versus-error determination, so this is not a no-review structure. What starts closed is the internal fairness check of the score itself.

  • assumed

    The retraining loop — past control outcomes feeding the next score alongside the annual random-audit survey — is drawn present, following the public critique that historically over-controlled groups feed forward into future high scores. Marked assumed, not baseline: the verified methodology documents the training labels as the random-audit survey, a deliberately chance-selected arm, so the feed-forward reading is a fair account of the critique rather than a documented mechanism of the training design.

  • baseline

    The household-data node carries a real inflow into the model — declared income and benefit amounts, household composition and life events, and behavioural/contact signals — so it enters the dynamics, not just the picture. That markers of economic vulnerability (low income, disability allowance, single parenthood) raise the score is documented on the model's own extracted arithmetic; what those signals encode as differential harm stays external (below).

  • baseline

    The defining feature is an absence: no independent fairness check ran on the score before outside scrutiny arrived, so the model-side check starts closed at baseline. The disproportion was provable on the model's own arithmetic and was later corroborated by the operator's own internal DSER simulation.

  • assumed

    The ranked control list is drawn as a mediating artifact on the model → controllers pathway — reflecting the case file's account of a score whose output prioritises which files to control. It carries no flow of its own and does not affect the dynamics.

  • assumed

    Documented disproportion in who this kind of model scores highest — RSA recipients, single mothers, foreign nationals — is recorded externally in the case file, hedged as the sources hedge (an internal simulation reported by journalists; a presumption of indirect discrimination found by the ombudsperson; disputed by CNAF; no court ruling). This Lab models institutional propagation, not demographics, and estimates no differential harm to served people.

What this example does not show

  • The documented harm is a contested disproportion in who the score ranks highest — RSA recipients, single mothers, and foreign nationals. The strongest figures come from CNAF's own internal DSER simulation reported by journalists, the ombudsperson found a presumption of indirect discrimination, CNAF disputes the framing, and no court had ruled. The Lab models institutional workflow propagation, not demographics, and estimates no differential harm to served people; that disproportion is documented in the case file and measured outside any diagram like this one.
  • False-positive rates by protected group were never released, and the redesigned 2025 model's real-world impact is so far simulated, not measured in production; the Lab uses the case's shape, not calibrated rates.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In an internal simulation study by CNAF's own statistics department (DSER), reported in October 2025 by Le Monde and La Quadrature du Net, recipients of the RSA minimum-income benefit were about 13% of beneficiaries but 39 to 41% of the highest-scoring 5%, and single mothers were about 14% of beneficiaries but 37 to 40% of that top bracket; households including a foreign national scored higher on average even after the nationality variable was removed. The full study is not public, and false-positive rates by protected group have not been released. The French ombudsperson (Défenseur des droits) told the Conseil d'État that a presumption of indirect discrimination appeared established because the differential treatment rests on beneficiaries' economic vulnerability; CNAF disputed the characterisation, and no court had ruled.

    empirical
    • Advocacy La Quadrature du Net, Notation des allocataires : la CNAF publie son code mais omet l'essentiel (Scoring of beneficiaries: CNAF publishes its code but omits the essential) (2026) https://www.laquadrature.net/2026/02/26/notation-des-allocataires-la-cnaf-publie-son-code-mais-omet-lessentiel/
    • Trade press Generation-NT, L'algorithme de la CAF est desormais dans le viseur de 25 organisations et du Defenseur des droits (The CAF algorithm is now in the sights of 25 organisations and the ombudsperson) (2026) https://www.generation-nt.com/actualites/caf-algorithme-discrimination-recours-conseil-etat-2069598
    • Advocacy La Quadrature du Net, CNAF's discriminatory scoring algorithm: 10 new organisations join the case before the Conseil d'Etat (2026) https://www.laquadrature.net/en/2026/01/20/cnafs-discriminatory-scoring-algorithm-10-new-organisations-join-the-case-before-the-conseil-detat-in-france/
  • France's family-benefits fund (CNAF) computes a monthly benefit-fraud suspicion score, on a 0-to-1 scale, for every benefit-receiving household — analysing the data of about 32 million people and producing more than 13 million scores each month, close to half of France's population; the highest scores route households into fraud controls, up to the most invasive on-site checks. An analysis by Le Monde and Lighthouse Reports of an extracted production model (a logistic regression of about 33 variables) found that markers of economic vulnerability raised the score: a stable-income family averaged about 0.33, while a person working while receiving the disability allowance (AAH) averaged about 0.66. The model's target was an overpayment (indu) above a threshold, which is frequently unintentional administrative error rather than proven intentional fraud, and the score itself is not disclosed to the person and cannot be appealed directly. CNAF disputed the discrimination framing, describing the tool as a neutral decision-aid that only prioritises which files to check; a coalition that grew to 25 organisations challenged the model before the Conseil d'État, and as of this writing no court had ruled.

    empirical
    • Investigative Lighthouse Reports, How We Investigated France's Mass Profiling Machine (methodology) (2023) https://www.lighthousereports.com/methodology/how-we-investigated-frances-mass-profiling-machine/
    • Investigative Lighthouse Reports, France's Digital Inquisition (2023) https://www.lighthousereports.com/investigation/frances-digital-inquisition/
    • Advocacy La Quadrature du Net, Scoring of welfare beneficiaries: the indecency of CAF's algorithm now undeniable (2023) https://www.laquadrature.net/en/2023/11/27/scoring-of-welfare-beneficiaries-the-indecency-of-cafs-algorithm-now-undeniable/
    • Advocacy La Quadrature du Net, CNAF's discriminatory scoring algorithm: 10 new organisations join the case before the Conseil d'Etat (2026) https://www.laquadrature.net/en/2026/01/20/cnafs-discriminatory-scoring-algorithm-10-new-organisations-join-the-case-before-the-conseil-detat-in-france/
    • Advocacy Amnesty International, France: Discriminatory algorithm used by the social security agency must be stopped (2024) https://www.amnesty.org/en/latest/news/2024/10/france-discriminatory-algorithm-used-by-the-social-security-agency-must-be-stopped/
    • Trade press Generation-NT, L'algorithme de la CAF est desormais dans le viseur de 25 organisations et du Defenseur des droits (The CAF algorithm is now in the sights of 25 organisations and the ombudsperson) (2026) https://www.generation-nt.com/actualites/caf-algorithme-discrimination-recours-conseil-etat-2069598

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Public benefits & eligibility domain page.

Levers available here and the patterns behind them

Documented case histories