PAN Lab example
INSS automated benefit analysis
Automation as queue management: when the metric makes denial the fastest way out
A benefit backlog in the millions, and a productivity metric that counts processes analyzed rather than the quality of each decision. Under that metric denial is the fastest way to clear a case, so automation used to drain the queue tilts toward the fastest disposition, not the correct one, and a fast denial is corrected only after a slow judicial channel that runs for years. Modeled on Brazil's INSS automated benefit-analysis record. The engine is not a fraud score and is not, in the ordinary sense, inaccurate; the leverage is a live merit check at the point of decision and a governance move that breaks the metric's grip, not a better engine. What an external audit measured, it measured after the denials had already persisted.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Auto-analysis-class benefit system driven by a case-count productivity metric network: 8 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 7 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the queue-management automation pattern documented in the Brazil INSS automated benefit-analysis case file — not a reconstruction of the actual INSS or Dataprev systems.
- baseline
The operator self-loop and the low correction rate encode the audited perverse incentive: INSS measures server productivity by the number of processes analyzed rather than the quality of the decision, so denial becomes the fastest disposition. In published PAN runs the same model embedded in different office cultures produces very different steady-state error adoption, so it is the culture the productivity metric creates — not the engine's accuracy — that drives the harm here.
- assumed
The engine is drawn as a rules-based administrative concession-and-denial engine (Track A) coupled to a documentary-conformity analysis by a human medical examiner (Track B, Atestmed), not a predictive or machine-learning risk score. The 10.94 percent automatic and 13.20 percent manual figures are TCU audit nonconformity (desconformidade) rates from samples — an audit-analysis outcome category that includes but is not identical to a court-confirmed wrongful-denial rate — not per-interaction generation rates.
- assumed
The benefit-determination record is drawn with slow audit-and-correct and a latent at-determination check to encode a two-speed error correction: a fast denial is reversed only through a slow judicial channel, with pending previdenciario cases averaging about 746 days, while the 30-day recurso is a faster but partial administrative channel. The error therefore persists as harm for years.
- assumed
The TCU audit is drawn as real but lagged oversight: it measured the nonconformity and ordered fixes (including a 180-day deadline to change the automatic-concession system so the insured are notified of discrepancies), but its at-determination reconciliation check starts inactive because the audit samples decisions after they have already persisted. The gap the case turns on is a live merit check at the point of decision, not the absence of oversight.
- assumed
The backlog-queue node is a mediating artifact naming the queue-management signature (automation used to drain a pending-request backlog under a case-count metric). It carries no flow of its own and does not affect the dynamics.
- assumed
The digital-intake feed is marked privacy-sensitive because the engine cross-checks CNIS contribution records and uploaded medical certificates to decide automatically. The cross-check is a legitimate part of adjudication; the privacy lever here governs which sensitive sources feed the automatic decision, not whether the decision runs.
- assumed
A digital-only front door excludes vulnerable applicants: a 2024 functional-literacy index found about 48 percent of Brazilians aged 50 to 64 performed poorly on digital-competency tests, so some never complete the digital intake at all — a denial-by-attrition that never enters the denial statistics. That differential exposure is documented in the case file and recorded outside any diagram like this one; this Lab models institutional propagation, not demographics, and estimates no differential harm to served claimants.
What this example does not show
- The 10.94 percent (automatic, January to May 2024) and 13.20 percent (manual, 2023) figures are TCU audit nonconformity (desconformidade) rates from samples, an audit-analysis category that includes but is not identical to a court-confirmed wrongful-denial rate; they are not a hard error rate. The absolute counts reported in coverage (about 920,000 automatic denials in the audited window, about 100,000 estimated wrongful, and 250,000 to 290,000 unjustified manual denials) are journalistic extrapolations from the TCU percentages, not officially published counts.
- The Lab models institutional workflow propagation, not people. The roughly 48 percent of applicants aged 50 to 64 with low digital competency who never complete the digital-only intake are a denial-by-attrition that never enters the denial statistics; that exclusion is documented in the case file and measured outside any diagram like this one, and the Lab estimates no differential harm to served claimants.
- There is no machine-learning risk-scoring model in the documented record. Automatic here means rules-based administrative processing (Track A) and documentary-conformity analysis by a human medical examiner (Track B, Atestmed); the two tracks are distinct, and this scenario models the queue-management incentive, not a predictive score.
- The 5.1 million pending previdenciario lawsuits and the roughly 746-day pending-case duration are CNJ caseload figures and cannot be mechanically attributed to automated denials specifically; the public data do not link an individual court reversal to the channel (automatic, manual, or Atestmed) that produced the denial.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In the sociotechnical simulation, the same AI in three modeled office cultures - stylized, not real workplaces - let errors stick at very different rates: roughly 75% under low-oversight autonomy, 20% under human supervision, and 16% under high-governance professional controls.
scenarioillustrative PAN-run resultNo published source is attached to this claim yet.
Where this connects
Institutional pressures in this domain
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Public benefits & eligibility domain page.
Levers available here and the patterns behind them
- Assign a challenger — Structured dissent
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Pause AI on alarms — Deployment circuit-breaker
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Store less data — Data minimization
- Upgrade model — Improve the model
- Escalate checks — State-feedback vigilance
Documented case histories
- INSS auto-analysis: when the productivity metric makes denial the fastest way out
- Michigan MiDAS
- Robodebt (Australia)
- Indiana / IBM eligibility modernization
- Rotterdam welfare-fraud risk model
- Arkansas ARChoices / ARIA
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- SyRI (Netherlands)
- CNAF benefit-fraud risk score (France)
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Udbetaling Danmark data-driven control (Denmark)
- BOSCO (Spain)
- Serbia Social Card (Socijalna karta)
- UK DWP Universal Credit Advances fraud model
- ID.me identity verification as an unemployment eligibility gate
- Medicaid unwinding: automated ex parte renewal at population scale
- Samagra Vedika
- Workforce Australia Targeted Compliance Framework: automated payment sanctioning after Robodebt
- NYC MyCity business chatbot
- Nevada DETR generative-AI unemployment appeals
- Tennessee TennCare TEDS