PAN Lab example
INSS automated benefit analysis
Automation as queue management: when the metric makes denial the fastest way out
Brazil's social security institute, INSS, uses automation to clear a benefit backlog. It counts staff output in cases analyzed, and denial finishes a case fastest.
See more
Brazil's social security institute, Instituto Nacional do Seguro Social (INSS), decides benefits in two tracks behind one app, Meu INSS. In one, a rules-based engine checks a person's social security contribution history, then grants, denies, or routes the request to manual analysis. In Atestmed, federal medical examiners decide temporary-incapacity benefits from uploaded certificates, without an in-person exam.
What the engine is and is not
Dataprev builds and runs the INSS processing systems. The engine runs on INSS and Dataprev systems.
The case file calls both tracks automated, but in Atestmed people decide. There, the automated part is the shared app intake. The examiners review documents for conformity, and the sources do not say against what.
Neither track is a machine-learning or predictive risk score. The case file says the engine is not a fraud score. It is not accused of illegality, and it is not, in the ordinary sense, inaccurate.
What the audit found
The TCU, the Tribunal de Contas da Uniao, is Brazil's federal audit court. It sorts audited denials into conforming and nonconforming. Nonconformity, desconformidade in Portuguese, includes wrongful denials, but it is not the same as a denial a court confirmed was wrong.
The TCU's plenary ruling on its audit, Acordao 634/2025-Plenario, came on 26 March 2025. It found manual denials 13.20 percent nonconforming in a 2023 sample. It found automatic denials 10.94 percent nonconforming from January to May 2024. Both were above the maximum acceptable limit.
The cause the audit named
INSS measures staff productivity by the number of processes analyzed, not the quality of each decision's justification. So denial is the fastest way to finish a case. The TCU found no incentive to justify a denial correctly. It found no effective communication with the insured.
Automation used to clear the backlog takes on that metric. It tilts toward the fastest decision, not the correct one.
What press coverage estimated
Legal-press coverage extrapolated counts from the TCU percentages. It estimated INSS granted about 5.964 million benefits in 2023. It estimated roughly 920,000 automatic denials in the audited window, about 100,000 of them wrongful. It also estimated 250,000 to 290,000 unjustified manual denials.
These are journalistic estimates, not officially published figures.
How a wrong denial is corrected
A denial is fast to make, and its correction is slow. A person can first file a recurso, the 30-day administrative appeal. If that does not resolve it, the person can go to the federal courts.
The Conselho Nacional de Justica (CNJ), which tracks Brazil's courts, recorded 5,109,076 pending social security lawsuits as of 31 October 2024. Pending cases averaged about 746 days. So a wrong denial taken to court is reversed only after roughly two years.
The TCU audit measured the denials after they had already been recorded and had led to benefit cutoffs.
What the TCU asked for next
In a 2026 follow-up, the TCU turned to automatic grants that leave out discrepancy notices. People entitled to a higher benefit are silently underpaid. It gave INSS, Dataprev, and Brazil's social security ministry 180 days to change the system so the insured are notified.
Both automated tracks remain in use and are expanding under reforms the TCU ordered.
What the case file reads
The case file says the changes that matter most are to the metric, and a check of each denial's merits when it is made. A better engine is not among them.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it goes on.
This case has a budget of 12 units and offers eleven tools. Assign a challenger costs 2 units, or 3 at its stronger setting. Review on schedule and Escalate checks cost 2 each, or 4 each at their stronger settings. Check with a second model, Store less data, and Upgrade model cost 3 each, or 5 each at their stronger settings. Check copied records costs 3, or 4 at its stronger setting.
Gate record entries and Vet connections cost 3 each, and Pause AI on alarms costs 4. These three have no stronger setting. Understand the system costs 3 under Explore (No Targets) and Service Targets Only, and 4 under the two higher levels. Its stronger setting costs 6.
While Understand the system is on, four tools cost 1 unit less: Check copied records, Pause AI on alarms, Vet connections, and Store less data. While it is at its stronger setting, they cost 2 less instead, but never less than 1 unit.
Lingering effects is a Dynamics setting in which damage outlasts its cause. It is on by default, off under Service and Safety Targets, and always on under All Governance Targets. While it is on, Check copied records, Vet connections, and Store less data work at reduced strength unless Understand the system is on.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. With the default Dynamics setting, the cheapest ways cost 5 units. One is Gate record entries with Escalate checks. Another is Understand the system with Vet connections.
The other two 5-unit ways are Assign a challenger with Gate record entries, and Assign a challenger at its stronger setting with Escalate checks. With Dynamics off, Vet connections alone meets the targets, at 3 units. No combination that includes Pause AI on alarms meets them, because halting the engine's cases lowers the service the network delivers.
Under Service and Safety Targets, the targets can be met. This level asks you to close all seven failure pathways open at the start, among other targets. Five different sets of tools meet it. Every set includes Assign a challenger, Escalate checks, and Vet connections. Each also includes Gate record entries or Store less data. The cheapest sets cost 10 units.
Each of those tools closes different pathways. Vet connections closes the pathways named Applications and records checked by the engine, and Recorded denial takes effect. Escalate checks closes the two pathways from the engine to the analysts and to the Atestmed examiners. Assign a challenger closes the pathway named Deny-fast habit shared among analysts.
Gate record entries or Store less data closes the pathways named Engine decisions written to the record, and Analysts' decisions recorded. Pause AI on alarms closes the same two pathways as Escalate checks. It also halts the engine's cases, and the service target is missed.
Under All Governance Targets, one set of tools meets the targets, and it spends all 12 units. It is Understand the system, Assign a challenger, Vet connections, Store less data, and Escalate checks. Understand the system works at either setting. At its stronger setting, Vet connections and Store less data cost 1 unit each.
The case file names four governance moves. The first, a quality and peer challenge to a denial, is Assign a challenger. The other three are an independent merit re-check, a check of each denial against the person's record, and a review rhythm closer to each decision.
Their tools here are Check with a second model, Check copied records, and Review on schedule. On this network they close no failure pathway, and none of them is needed to meet any level. Upgrade model and Pause AI on alarms are in no combination that meets the two higher levels.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Auto-analysis-class benefit system driven by a case-count productivity metric network: 8 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 7 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 8 assumptions
- assumed
This example follows the pattern the INSS case file documents: automation used to clear a backlog of requests under a metric that counts cases analyzed. It is not a copy of the actual INSS or Dataprev systems.
- baseline
This example shows the incentive the TCU audited in two places. One is the pathway named Deny-fast habit shared among analysts. The other is the analysts' correction of mistakes, which the metric does not reward. INSS measures staff productivity by the number of processes analyzed, not the quality of the decision. So denial becomes the fastest way to finish a case. In the sociotechnical simulation behind the Lab, the same AI in different modeled office cultures let errors stick at very different rates. So this example treats the culture the metric creates, not the engine's accuracy, as what drives the harm.
- assumed
This example draws the engine as a rules-based system that grants or denies benefits, the first track. Beside it, in the second track, Atestmed, a human medical examiner reviews documents for conformity. Neither is a predictive or machine-learning risk score. The 10.94 percent automatic and 13.20 percent manual figures are TCU nonconformity rates from audit samples. Nonconformity, desconformidade in Portuguese, is an audit category. It includes wrongful denials, but it is not the same as a denial a court confirmed was wrong. These figures do not say how often the engine errs on any single case.
- assumed
This example draws two speeds of correction on the benefit-determination record. A denial is fast to make. Reversing it in court is slow: pending social security cases average about 746 days. The 30-day recurso is a faster administrative appeal, and the sources call it partial. The sources describe no check of a denial when it is recorded. So a wrong denial persists as harm for years.
- assumed
This example treats the TCU audit as real oversight that comes late. The audit measured the nonconformity and named its cause. In a 2026 follow-up, the TCU gave INSS, Dataprev, and Brazil's social security ministry 180 days to change the automatic-grant system. The change must notify the insured of discrepancies. The audit sampled decisions after they had been recorded. So what this case turns on is a check of each denial's merits when it is made, not whether any oversight exists.
- assumed
The part named Pending-request backlog stands for the pressure INSS answers to: clearing a backlog of requests under a metric that counts cases analyzed. No pathway connects to it, and it changes nothing about how mistakes pass.
- assumed
The pathway named Applications and records checked by the engine is marked as handling sensitive personal data. The engine decides from contribution records, and people upload medical certificates through the same channel. Checking them is a legitimate part of deciding a benefit. The privacy tools here govern which sensitive sources the automatic decision uses, not whether the decision runs.
- assumed
The digital-only channel shuts out some vulnerable applicants. A 2024 functional-literacy index, Inaf, found about 48 percent of Brazilians aged 50 to 64 performed poorly on digital-competency tests. The digital-only application stops some applicants before they finish it. Their exclusion never appears in the denial statistics. The case file documents it. This example traces how mistakes pass inside the institution, not who is affected. It estimates no difference in harm among the people served.
What this example does not show
Show all 4 limitations
- The 10.94 percent figure covers automatic denials from January to May 2024. The 13.20 percent figure covers manual denials in 2023. Both are TCU nonconformity rates from audit samples. Nonconformity, desconformidade in Portuguese, includes wrongful denials, but it is not the same as a denial a court confirmed was wrong. So neither is a hard error rate. Press coverage extrapolated counts from these percentages. It estimated about 920,000 automatic denials in the audited window, about 100,000 of them wrongful. It also estimated 250,000 to 290,000 unjustified manual denials. These are journalistic estimates, not officially published counts.
- This example traces how mistakes pass inside the institution, not what happens to people. A 2024 index found about 48 percent of Brazilians aged 50 to 64 performed poorly on digital-competency tests. The digital-only application stops some applicants before they finish it, so their exclusion never enters the denial statistics. The case file documents that exclusion, and this example does not show it. It estimates no difference in harm among the people served.
- The sources document no machine-learning risk-scoring model. Automatic here means two things. One is rules-based administrative processing, the first track. The other is documentary review by a human medical examiner, the second track, Atestmed. The two tracks are distinct. This example shows the incentive to clear the queue fast, not a predictive score.
- The 5.1 million pending social security lawsuits and the average of about 746 days per pending case are CNJ caseload figures. They cannot be tied to automated denials in particular. The public data do not link a court reversal to the channel that produced the denial: automatic, manual, or Atestmed.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In the sociotechnical simulation, the same AI in three modeled office cultures, stylized and not real workplaces, led to very different outcomes. Mistakes built on one another far more under low-oversight autonomy than under human supervision or high-governance professional controls.
scenarioillustrative PAN-run resultNo published source is attached to this claim yet.
Where this connects
Institutional pressures in this domain
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Public benefits & eligibility domain page.
Levers available here and the patterns behind them
- Assign a challenger — Structured dissent
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Pause AI on alarms — Deployment circuit-breaker
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Store less data — Data minimization
- Upgrade model — Improve the model
- Escalate checks — State-feedback vigilance
Documented case histories
- INSS auto-analysis: when the productivity metric makes denial the fastest way out
- Michigan MiDAS
- Robodebt (Australia)
- Indiana / IBM eligibility modernization
- Rotterdam welfare-fraud risk model
- Arkansas ARChoices / ARIA
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- SyRI (Netherlands)
- CNAF benefit-fraud risk score (France)
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Udbetaling Danmark data-driven control (Denmark)
- BOSCO (Spain)
- Serbia Social Card (Socijalna karta)
- UK DWP Universal Credit Advances fraud model
- ID.me identity verification as an unemployment eligibility gate
- Medicaid unwinding: automated ex parte renewal at population scale
- Samagra Vedika
- Workforce Australia Targeted Compliance Framework: automated payment sanctioning after Robodebt
- NYC MyCity business chatbot
- Nevada DETR generative-AI unemployment appeals
- Tennessee TennCare TEDS