Skip to content

PAN Lab example

Forsakringskassan VAB fraud-selection profile (Sweden)

The audit the agency refused: a secret fraud-selection profile no one outside could see

Försäkringskassan, Sweden's Social Insurance Agency, used a machine-learning profile to pick sick-child benefit claimants for fraud and error investigation. It never released the model.

See more

The risk-based selection profile was a machine-learning model that Försäkringskassan, the Swedish Social Insurance Agency, built and ran in house. It scored applicants for the temporary parental allowance, paid to parents caring for a sick child. The highest-scored went to the agency's control department for fraud and error investigation.

How it was used

Amnesty International reported that the system was in use since at least 2013. It was a targeting step, not an automatic denial. Human investigators in the control department worked each flagged case before any sanction.

In 2017, about 6,129 people were selected for investigation, out of roughly 977,730 applicants. The profile chose 5,082 of them, and 1,047 were picked at random.

The people selected were not told that an algorithm had flagged them. Their payments could be delayed while the investigation ran.

The random sample the agency held

The agency ran the random arm as a deliberate part of its design. Because those applicants were picked by chance, their results give an unbiased base rate of errors.

The agency also uses a random sample to estimate wrong payments, and case workers found at least one wrongly paid day in 20.2 percent of applications. The case workers did not establish whether a mistake was deliberate. The sources do not say whether that is the random arm.

The case file calls the random arm exactly the instrument an audit needs. The inspectorate ISF used the random arm in 2018 to test the profile. The agency said it ran its own equal-treatment comparison against the random arm, but refused to disclose the results.

What outside reviews found

In 2018 the inspectorate ISF published two reviews. It found the profiling substantially more accurate than other control methods, but raised legal-certainty and equal-treatment concerns. It concluded that the profile, in its current design, did not pass for equal treatment. The agency rejected that conclusion.

On 27 November 2024, Lighthouse Reports and the newspaper Svenska Dagbladet published "Sweden's Suspicion Machine." They obtained the agency's 2017 outcome data, but not the model.

By their analysis, the profile over-selected women, people of foreign background, below-median earners, and people without a university degree. It also wrongly flagged those groups at higher rates.

These are outcome computations under specific fairness definitions, from one year of data. The agency disputed the discrimination framing.

The refused disclosure

The agency never released the model, its input variables, or its precision. It resisted freedom-of-information requests for about three years, saying disclosure would help fraudsters. It refused to say whether its models were trained on random samples.

Amnesty International reported that in 2020 a former data protection officer at the agency warned that the operation broke European data-protection rules.

How it ended

In June 2025 the privacy regulator IMY opened a supervision case into whether the selection complied with the GDPR, the EU data-protection law. In a response dated 5 September 2025, the agency said it had taken the profile out of service about a month earlier.

IMY closed the case on 18 November 2025, because the system was no longer used. That closure is not a ruling that the system was lawful or free of discrimination. No court judgment and no regulatory fine were issued.

Where to look

The case file calls this the case of the refused audit. It reads past investigation outcomes as the profile's training labels, so the groups investigated most are selected again. The agency refused to say whether its models were trained on random samples.

In the case file's reading, the remedy has to come from forcing the audit the agency refused. It also has to come from checking the profile against the random sample the agency already held, not from a better score.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Here a mistake means a wrong selection or finding in the agency's work, not an error in an application. A pathway is closed when mistakes stop passing along it.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. They ask that mistakes stop building on one another, and that the profile stay useful to the investigators' work.

The cheapest combinations cost 4 of this case's 12 budget units. Each pairs two of three tools: Mark AI-written records, Escalate checks, and Peer sharing rules.

Under Service and Safety Targets and under All Governance Targets, the targets can also be met. Both levels ask you to close every failure pathway, among other targets. Every combination that meets them includes four tools: Mark AI-written records, Escalate checks, Peer sharing rules, and Gate record entries. Those four cost 9 units and close all eight failure pathways.

Understand the system costs 3 units under Explore and Service Targets Only, and 4 under the two higher levels. While it is on, it cuts 1 unit from each of four tools: Check copied records, Escalate checks, Upgrade model, and Peer sharing rules.

The audit tools the case file stresses are Review on schedule, Require sign-off, Check copied records, and Assign a challenger. They add checks or slow how mistakes build up. None of them is needed to meet the targets here, and only Assign a challenger closes a pathway.

More tools are not better here. Every tool at once, ignoring the budget, misses the service target at every level that has one. Some tools take something from the service, and together they take it just below its target.

Stylized model of a documented deploymentPublic benefits & eligibility

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Forsakringskassan-class risk-based fraud-selection profile network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 6 assumptions
  • assumed

    This example follows the fraud-selection profile described in the Sweden Försäkringskassan case file. It does not rebuild the actual model, whose features were never disclosed. The disparity figures come from one obtained year of outcome data, tested against specific fairness definitions. They are not confirmed model internals. The agency disputed the discrimination framing.

  • assumed

    This example assumes mistakes pass sideways in two ways. The same profile scores each applicant it is run on, and investigators share suspicion patterns. Against both, it includes a standing outside review of who the profile keeps flagging. The deployment never had a standing review like it.

  • baseline

    This example assumes past investigation outcomes train the profile, so groups investigated more are selected again. That reading comes from the case file. The agency refused to say whether its models were trained on random samples. The example shows the shape of that loop, not a measured rate.

  • baseline

    The random-sample results were a real store the agency held. In 2017, 1,047 of the 6,129 investigations were picked at random. This example draws a check comparing the control records with those results. The agency confirmed it used an equal-treatment procedure making such a comparison, but refused to disclose the results.

  • baseline

    The agency withheld the model, its features, and its precision from outsiders. It refused information requests for about three years. So this example assumes no standing outside review of who the profile flagged. It assumes no routine, published check against the random sample either. No outside audit schedule or deployment sign-off was in force. The disclosure and audit tools supply those controls.

  • assumed

    The reported harm falls on particular groups. The profile over-selected women, people of foreign background, below-median earners, and people without a university degree. It also wrongly flagged those groups more often. This example shows how errors move through the agency's work, not who they fall on. It estimates no difference in harm between groups. The case file records those differences, measured outside any diagram like this one.

What this example does not show

Show all 2 limitations
  • This example does not show who bears the harm. The profile over-selected women, people of foreign background, below-median earners, and people without a university degree. It also wrongly flagged those groups at higher rates. The example shows how errors move through the agency's work, not demographics. The case file records those differences, measured outside any diagram like this one.
  • The disparity figures come from one obtained year of outcome data, tested against specific fairness definitions. They are not confirmed model internals. The agency never released the model, its features, or its precision, and it disputed the framing. No court or regulator ruled that the profile discriminated or breached the GDPR. The privacy regulator IMY closed its supervision because the system had been withdrawn. This example uses the case's shape, not calibrated rates.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Analysing the Swedish Social Insurance Agency (Forsakringskassan) 2017 outcome data, Lighthouse Reports and Svenska Dagbladet reported on 27 November 2024 that the agency's in-house machine-learning risk profile for the temporary parental allowance (VAB) selected women (more than 1.5x), people of a foreign background (about 2.5x), below-median earners (2.97x), and people without a university degree (3.31x) for fraud investigation more often than comparison groups by demographic parity, and wrongly flagged those groups at higher false-positive rates (about 1.7x for women and 2.4x for people of a foreign background); in a random sample the agency uses to estimate wrong payments, case workers found at least one incorrectly paid day in 20.2 percent of applications, without establishing intent. These are outcome computations under specific fairness definitions from a single obtained year of data, not confirmed model internals; the agency disputed the framing and did not release the model. The data-protection regulator IMY closed its GDPR supervision on 18 November 2025 for mootness after the agency withdrew the system, and no court or regulator issued a discrimination or GDPR penalty.

    empirical
    • Investigative Lighthouse Reports, Sweden's Suspicion Machine (co-published with Svenska Dagbladet, 27 Nov 2024) https://www.lighthousereports.com/investigation/swedens-suspicion-machine/
    • Investigative Lighthouse Reports, How we investigated Sweden's Suspicion Machine (methodology) https://www.lighthousereports.com/methodology/sweden-ai-methodology/
    • Investigative Lighthouse Reports, suspicion_machines_sweden (data and analysis repository, GitHub) https://github.com/Lighthouse-Reports/suspicion_machines_sweden
    • Government Integritetsskyddsmyndigheten (IMY), Avslutad tillsyn efter att Forsakringskassan tagit AI-system ur bruk (Supervision closed after Forsakringskassan took AI system out of use) (2025) [Swedish] https://www.imy.se/nyheter/avslutad-tillsyn-efter-att-forsakringskassan-tagit-ai-system-ur-bruk/
    • Government Integritetsskyddsmyndigheten (IMY), Tillsyn: Forsakringskassan (supervision case page and decision, 18 Nov 2025) (2025) [Swedish] https://www.imy.se/tillsyner/forsakringskassan/
  • The Swedish Social Insurance Agency (Forsakringskassan) did not disclose the machine-learning risk profile it used to select temporary-parental-allowance recipients for fraud investigation: its algorithm class, features, and precision were never released, and the agency resisted freedom-of-information disclosure for roughly three years on fraud-prevention grounds. In 2018 the audit inspectorate ISF found the risk-based profiling substantially more accurate than alternative controls while warning that it raised legal-certainty and equal-treatment concerns, and cautioning that an accurate model can still be inequitable when two groups err equally but only one is followed up. Amnesty International reported that a former agency data protection officer warned in 2020 that the operation breached European data-protection rules. The system was decommissioned in 2025 during the regulator's supervision, before any court or regulator ruled on it.

    empirical
    • Investigative Lighthouse Reports, Sweden's Suspicion Machine (co-published with Svenska Dagbladet, 27 Nov 2024) https://www.lighthousereports.com/investigation/swedens-suspicion-machine/
    • Investigative Lighthouse Reports, How we investigated Sweden's Suspicion Machine (methodology) https://www.lighthousereports.com/methodology/sweden-ai-methodology/
    • Government evaluation Inspektionen for socialforsakringen (ISF), Profilering som urvalsmetod for riktade kontroller (Profiling as a selection method for targeted controls) (2018) [Swedish] https://isf.se/publikationer/rapporter/2018/2018-03-26-profilering-som-urvalsmetod-for-riktade-kontroller
    • Government evaluation Inspektionen for socialforsakringen (ISF), Riskbaserade urvalsprofiler och likabehandling (Risk-based selection profiles and equal treatment) (2018) [Swedish] https://isf.se/publikationer/rapporter/2018/2018-06-15-riskbaserade-urvalsprofiler-och-likabehandling
    • Advocacy Amnesty International, Sweden: Authorities must discontinue discriminatory AI systems used by welfare agency (2024) https://www.amnesty.org/en/latest/news/2024/11/sweden-authorities-must-discontinue-discriminatory-ai-systems-used-by-welfare-agency/

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Public benefits & eligibility domain page.

Levers available here and the patterns behind them

Documented case histories