Skip to content

PAN Lab example

CNAF benefit-fraud risk score (France)

The score that suspects the vulnerable: a benefit-fraud risk model

France's family-benefits fund, CNAF, scores every benefit-receiving household for fraud risk each month. In the versions examined, markers of economic vulnerability raised the score.

See more

The CNAF risk score, known as the “datamining” score, is software that France's family-benefits fund built and runs in house. Each month it gives every benefit-receiving household a fraud-suspicion score between 0 and 1. Controllers at the family-allowance funds open fraud controls on the highest scores first.

What the score weighs

The score comes from a logistic regression, a statistical model that combines weighted facts about a household into one number. The 2014 version weighed about 33 such facts. It covers the data of about 32 million people, close to half of France's population. That makes roughly 13 million scores each month.

It reads what the fund holds on each household: declared income and benefit amounts, household makeup, life events, and contact habits, such as online log-ins.

What raises the score

Le Monde and Lighthouse Reports analysed the 2014 version. On its own weights, markers of economic vulnerability raised the score. They included low income, unemployment, receiving the RSA minimum-income benefit, and a high rent compared with income. Others were rarely logging in online, a recent separation or move, errors in declarations, and being a single parent.

Before 2025, working while receiving the disability allowance, called the AAH, raised it too. Critics called that “the height of cynicism.” In the analysis, a family with a stable income averaged about 0.33. A person working while receiving the AAH averaged about 0.66.

What the score looks for

The score looks for overpayments (indus in French), not proven fraud. It was trained to find overpayments of at least 600 euros a month lasting at least six months. CNAF's own training manual says this target covers 98 percent of fraudulent payments. An overpayment is often an honest administrative or declaration error, not deliberate fraud.

CNAF's annual random-audit survey supplied the examples the score learned from. Each year the fund checks a random sample of files, whatever their score.

Who decides

CNAF calls the score a decision-aid that only sets which files are checked first. Controllers, not the score, decide whether a household was overpaid, and if so whether by fraud or error. The sources read do not say who decides the outcome of the roughly 29 million automated controls.

The most invasive controls happen on site. There, controllers can question who lives in the household, examine bank records, and visit the home.

The household never sees its score and cannot appeal it directly. It can contest only the control decision or the repayment demand that follows.

How many controls there are

Official figures for 2024 record 31.5 million controls involving 6.4 million benefit recipients. About 29 million were automated. The other 2.5 million were document checks or on-site visits. The national anti-fraud service's share of detected fraud rose from 48 million euros in 2021 to 166 million euros in 2024.

What the reporting found

Reporting found that since 2019 the score flagged more than 700,000 investigations. About 35 percent of flagged households ultimately had to repay, about 821 euros on average. About 17 percent were found to be owed money by CNAF. In random-audit controls, about 15 percent of households had to repay.

Of the more than 100,000 people sent each year to in-depth investigations, nearly seven in ten were chosen because of a high score. The analysis could not measure how often people in protected groups were flagged wrongly. Outcome data by protected group was not available.

What each side says

CNAF's own statistics department, the DSER, ran a simulation study. Le Monde and La Quadrature du Net reported it in October 2025. Recipients of the RSA minimum-income benefit were about 13 percent of beneficiaries, but 39 to 41 percent of the highest-scoring 5 percent. Single mothers were about 14 percent of beneficiaries, but 37 to 40 percent of that top group. Households including a foreign national scored higher, even after the nationality variable was removed. The full study is not public.

In October 2025, the French ombudsperson, the Défenseur des droits, filed observations with the Conseil d'État. It found that “a presumption of indirect discrimination appears established,” because the difference rests on beneficiaries' economic vulnerability. It said the tool appears to over-control the most precarious people.

CNAF disputes the discrimination framing. It calls the tool a neutral decision-aid that targets significant and repeated overpayments. One CNAF director argued it was “the opposite of discrimination,” because no one can explain why a given file is targeted. No court had ruled as of the sources read.

How the score was opened up

In 2023, Le Monde and Lighthouse Reports obtained earlier versions of the model under France's freedom-of-information law. They went through the CADA, the Commission for Access to Administrative Documents, after CNAF resisted.

On 15 and 16 October 2024, fifteen organisations led by La Quadrature du Net asked the Conseil d'État to strike the score down. They cited data protection and non-discrimination. Amnesty International France was among them. In January 2026, ten more organisations joined, bringing the coalition to 25. A public hearing was expected in spring 2026.

The 2018 version challenged in the case ran until January 2026, and it had been withheld from the 2023 investigation. CNAF published the redesigned model's source code on 15 January 2026. La Quadrature du Net reported in late February 2026 that CNAF had just quietly published the 2018 version. Its headline said CNAF published its code but left out the essentials.

The redesigned model

A redesigned “2025” model entered production in January 2026. It removed working while receiving the AAH, nationality, housing type, and behaviour as variables.

Critics, including La Quadrature du Net, say it still over-targets the same groups. It keeps triggers such as receiving three or more benefits, receiving over 200 euros a month in benefits, low income, and starting or leaving the RSA. Its effect so far comes from the DSER's simulations, not from measuring it in production.

What this case asks

Watch two things at once. First, the facts that raise the score are the markers of the people a safety net exists to serve. The case file calls this a double penalty: the people the safety net serves are the people it suspects. Counting ordinary errors as overpayments sharpens it.

Second, watch for a loop. CNAF's documented training method uses the random-audit survey. If the results of past controls also shaped retraining, a skew in who gets checked would shape who the score suspects next. The sources do not document that step.

The case also shows that opening a score up is not the same as stopping it. Its weights were forced into view, and the system kept running.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A pathway counts as closed once it passes on only a few mistakes. It need not stop them all.

This case's budget is 13 units. Each tool costs units from it.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest combinations use two tools and cost 4 of the 13 units. Each pairs Mark AI-written records with Escalate checks, Peer sharing rules, or Assign a challenger.

Under Service and Safety Targets and under All Governance Targets, the targets can also be met. Both levels ask you to close every failure pathway, among other targets. The cheapest combinations cost 11 of the 13 units. They use four tools: Store less data, Mark AI-written records, Peer sharing rules, and Escalate checks at its stronger setting.

Every combination that meets these targets includes Store less data and Escalate checks at its stronger setting. Escalate checks at that setting is the one tool that closes High scores sent to controls. Store less data is the one tool that closes Outcomes written to records and Scores logged to records.

Understand the system costs 4 units at these two levels and 3 at the lower two. While it is in use, it cuts the price of Escalate checks, Peer sharing rules, Mark AI-written records, and Store less data. The cut is 1 unit each. At its stronger setting, which costs 6 units, the cut is 2, though no tool drops below 1 unit. Adding it to the four tools above keeps the total at 11 units at either setting. At the standard setting, its 4 units are offset by 4 units of cuts. At the stronger setting, its 6 units are offset by 6, because Mark AI-written records and Peer sharing rules stop at 1 unit.

Three tools change no pathway on this network: Review on schedule, Keep skills sharp, and Upgrade model. The first limits how far mistakes build on one another. The other two act on the controllers and on the score itself.

Stylized model of a documented deploymentPublic benefits & eligibility

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the CNAF-class benefit-fraud risk-scoring model network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 6 assumed · 4 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 10 assumptions
  • baseline

    This example includes the outside bodies that examined the score: the CADA, the Défenseur des droits, and the Conseil d'État. Reporters obtained earlier versions of the model under freedom-of-information law. CNAF published the redesigned model's code in January 2026 and the 2018 version by late February. That disclosure took years of proceedings, not routine practice. A coalition that grew to 25 organisations filed a case to annul the score in October 2024. The ombudsperson filed observations in October 2025. The scrutiny was aimed at the model, and a redesigned version entered production in January 2026. It dropped working while receiving the disability allowance, nationality, housing type, and behaviour as variables. The scrutiny never became a scheduled review of who is selected, and no court had ruled.

  • baseline

    This example assumes the score's ranking picks most people sent to in-depth investigations. Of the more than 100,000 people sent each year to in-depth investigations, nearly seven in ten were chosen because of a high score. It also assumes the random-audit survey is the main source of the score's training examples, as CNAF's documented training method says. That method labels each example by a target that CNAF's training manual says covers 98 percent of fraudulent payments. Controllers' findings from targeted controls are assumed to play a smaller part. That part is a reading of the sources, not a documented step.

  • assumed

    The sources name a second input, so this example includes it: the annual random-audit survey. It supplies the score's training examples and the share of random checks that find an overpayment. It is the one part of this process that looks at households picked by chance, not by the score's own history. That makes it a fair benchmark, and it sets it apart from the household data the score reads.

  • assumed

    This example follows the national benefit-fraud scoring pattern documented in the case file on France's CNAF score. It does not rebuild the real model, its variables, or their weights.

  • assumed

    Between controllers, this example assumes two things. Controllers pass targeting habits to one another, which can spread mistakes. They also decide whether a household was overpaid, and whether by fraud or error, which is a real human check on the score. The sources show no independent fairness check of the score itself before outside scrutiny.

  • assumed

    This example assumes the score is retrained partly on past control outcomes, alongside the random-audit survey. That follows the case file's warning that, if so, groups checked more in the past would be scored higher later. It is marked as an assumption, because CNAF's documented training method uses the chance-picked survey. The sources do not document retraining on past controls.

  • baseline

    In this example the household data is a real input to the score: declared income and benefit amounts, household makeup, life events, and contact habits. The model's own weights show that markers of economic vulnerability raised the score. They include low income, working while receiving the disability allowance, and single parenthood. What those markers mean as unequal harm to people stays outside this example, as the last assumption explains.

  • baseline

    The sources show that no independent fairness check ran on the score before outside scrutiny arrived. This example includes that check so you can add it. The imbalance could be shown from the model's own weights. CNAF's internal DSER simulation later confirmed it.

  • assumed

    This example places the ranked control list between the score and the controllers. The case file describes a score whose output sets which files are checked first. The list has no pathways of its own, so it changes nothing in how mistakes move.

  • assumed

    The case file records a disputed imbalance in who this kind of score ranks highest: RSA recipients, single mothers, and foreign nationals. Its sources are hedged. A CNAF internal simulation was reported by journalists, and the ombudsperson found a presumption of indirect discrimination. CNAF disputes it, and no court had ruled. This example shows how errors move among the score, the controllers, and the records. It does not model groups of people or estimate unequal harm to them.

What this example does not show

Show all 2 limitations
  • The documented harm is a disputed imbalance in who the score ranks highest: RSA recipients, single mothers, and foreign nationals. The strongest figures come from CNAF's own internal DSER simulation, reported by journalists. The ombudsperson found a presumption of indirect discrimination, CNAF disputes the framing, and no court had ruled. This example shows how errors move among the score, the controllers, and the records. It does not model groups of people or estimate unequal harm to them. The case file documents that imbalance, which is measured outside any network like this one.
  • How often people in protected groups were flagged wrongly was never released. The redesigned 2025 model's real-world effect has so far been simulated, not measured in production. This example follows the shape of the case, not measured rates.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In an internal simulation study by CNAF's own statistics department (DSER), reported in October 2025 by Le Monde and La Quadrature du Net, recipients of the RSA minimum-income benefit were about 13% of beneficiaries but 39 to 41% of the highest-scoring 5%, and single mothers were about 14% of beneficiaries but 37 to 40% of that top bracket; households including a foreign national scored higher on average even after the nationality variable was removed. The full study is not public, and false-positive rates by protected group have not been released. The French ombudsperson (Défenseur des droits) told the Conseil d'État that a presumption of indirect discrimination appeared established because the differential treatment rests on beneficiaries' economic vulnerability; CNAF disputed the characterisation, and no court had ruled.

    empirical
    • Advocacy La Quadrature du Net, Notation des allocataires : la CNAF publie son code mais omet l'essentiel (Scoring of beneficiaries: CNAF publishes its code but omits the essential) (2026) https://www.laquadrature.net/2026/02/26/notation-des-allocataires-la-cnaf-publie-son-code-mais-omet-lessentiel/
    • Trade press Generation-NT, L'algorithme de la CAF est desormais dans le viseur de 25 organisations et du Defenseur des droits (The CAF algorithm is now in the sights of 25 organisations and the ombudsperson) (2026) https://www.generation-nt.com/actualites/caf-algorithme-discrimination-recours-conseil-etat-2069598
    • Advocacy La Quadrature du Net, CNAF's discriminatory scoring algorithm: 10 new organisations join the case before the Conseil d'Etat (2026) https://www.laquadrature.net/en/2026/01/20/cnafs-discriminatory-scoring-algorithm-10-new-organisations-join-the-case-before-the-conseil-detat-in-france/
  • France's family-benefits fund (CNAF) computes a monthly benefit-fraud suspicion score, on a 0-to-1 scale, for every benefit-receiving household — analysing the data of about 32 million people, close to half of France's population, and producing more than 13 million scores each month; the highest scores route households into fraud controls, up to the most invasive on-site checks. An analysis by Le Monde and Lighthouse Reports of an extracted production model (a logistic regression of about 33 variables) found that markers of economic vulnerability raised the score: a stable-income family averaged about 0.33, while a person working while receiving the disability allowance (AAH) averaged about 0.66. The model's target was an overpayment (indu) above a threshold, which is frequently unintentional administrative error rather than proven intentional fraud, and the score itself is not disclosed to the person and cannot be appealed directly. CNAF disputed the discrimination framing, describing the tool as a neutral decision-aid that only prioritises which files to check; a coalition that grew to 25 organisations challenged the model before the Conseil d'État, and as of this writing no court had ruled.

    empirical
    • Investigative Lighthouse Reports, How We Investigated France's Mass Profiling Machine (methodology) (2023) https://www.lighthousereports.com/methodology/how-we-investigated-frances-mass-profiling-machine/
    • Investigative Lighthouse Reports, France's Digital Inquisition (2023) https://www.lighthousereports.com/investigation/frances-digital-inquisition/
    • Advocacy La Quadrature du Net, Scoring of welfare beneficiaries: the indecency of CAF's algorithm now undeniable (2023) https://www.laquadrature.net/en/2023/11/27/scoring-of-welfare-beneficiaries-the-indecency-of-cafs-algorithm-now-undeniable/
    • Advocacy La Quadrature du Net, CNAF's discriminatory scoring algorithm: 10 new organisations join the case before the Conseil d'Etat (2026) https://www.laquadrature.net/en/2026/01/20/cnafs-discriminatory-scoring-algorithm-10-new-organisations-join-the-case-before-the-conseil-detat-in-france/
    • Advocacy Amnesty International, France: Discriminatory algorithm used by the social security agency must be stopped (2024) https://www.amnesty.org/en/latest/news/2024/10/france-discriminatory-algorithm-used-by-the-social-security-agency-must-be-stopped/
    • Trade press Generation-NT, L'algorithme de la CAF est desormais dans le viseur de 25 organisations et du Defenseur des droits (The CAF algorithm is now in the sights of 25 organisations and the ombudsperson) (2026) https://www.generation-nt.com/actualites/caf-algorithme-discrimination-recours-conseil-etat-2069598

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Public benefits & eligibility domain page.

Levers available here and the patterns behind them

Documented case histories