Skip to content

PAN Lab example

SSA Insight

The verifier and the sign flip: a decision-checking copilot

This copilot runs the other way. A human writes the disability decision first; the tool reads the draft and raises quality flags before it issues, and every fully favorable draft has to be run through it. Modeled on the Social Security Administration's Insight quality-flagging tool. Nothing here decides anything, and that is the point: it is a verifier. The trap is the sign flip - under caseload the check that catches errors becomes a checklist to satisfy, and the errors the flags never learned to see ride straight through wearing a corrected, human-signed decision.

Stylized model of a documented deploymentCaseworker documentation & copilots

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the SSA-Insight-class reversed-direction verifier copilot network: 7 components and 16 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the reversed-direction quality-flag-verifier pattern documented in the Social Security Administration (SSA) Insight case file — a copilot that checks human-drafted decisions rather than drafting them — not a reconstruction of the actual software.

  • assumed

    Peer pathways are authored on both signs: a pass-the-flags shortcut and a single verifier's systematic blind spots reinforce, the colleague cross-check survives at a low level, and the policy-office review of the flags starts closed — the gap the 2019 Inspector General audit named.

  • assumed

    The verifier reads the drafted decision and joins structured case data at baseline and writes nothing into the record; corrected drafts enter under a human signature, so whether error reaches the record depends on the flag-check habit holding, not on an unsupervised write.

  • assumed

    The quality-assurance (QA) output store is generated by every run and used for management reporting; whether it is reused to feed the verifier or further models is not established in the record, so that pathway is drawn inactive rather than active.

  • assumed

    Served disability claimants, and the benefits or remands they do or do not receive, are not in the dynamics. The tool uses personally identifiable information (PII) but no demographic features and no differential-harm audit of it exists, so no differential client harm is estimated here; those outcomes are documented in the case file and measured outside any diagram like this one.

What this example does not show

  • Served disability claimants, and the benefits or remands they do or do not receive, are not modeled here; the Lab models institutional propagation only, and those outcomes are documented in the case file and measured outside any diagram like this one.
  • The detection-lift and quality-return figures (about 0.9 versus 0.7 errors logged per case, 12.6 versus 21.5 percent of cases returned, about 4.7 fewer days per case) are internal, non-randomized agency comparisons with self-selected voluntary users, reported through the 2019 Inspector General audit; the agency stopped tracking performance after five months and could not determine any effect on remands, so these are decision-support signals, not independent causal findings.
  • The 43-flag count is from the 2025 federal AI inventory; the run-before-issuance mandate for fully favorable decisions and the study figures are attested to 2018-2019 sources; this scenario does not weld one era's mandate to another era's flag count, and the tool is classical predictive machine learning, not a generative model.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The US Social Security Administration requires decision writers to run fully favorable disability decisions through its in-house Insight verifier before issuance, with narrow documented exceptions, and the 2025 federal AI inventory records the tool computing 43 quality flags. In the agency's internal five-month study of roughly 50,000 appeals-level cases, reported through the 2019 Inspector General audit, analysts who used Insight logged about 0.9 errors per case against 0.7 for non-users, saw processing time fall about 4.7 days per case, and had about 12.6 percent of their cases returned for quality issues against 21.5 percent for non-users. These are internal, non-randomized comparisons among self-selected voluntary users, and the same audit found the agency stopped tracking performance after the first five months and could not determine any effect on remands.

    empirical
    • Government evaluation US Social Security Administration, Office of the Inspector General, The Social Security Administration's Use of Insight Software to Identify Potential Anomalies in Hearing Decisions (A-12-18-50353) (2019) https://oig-files.ssa.gov/audits/full/A-12-18-50353.pdf
    • Government US Social Security Administration, SSA Individual AI Use Case Inventory 2025 (2026) https://www.ssa.gov/sites/default/files/2026-04/SSA-Individual-AI-Inventory-2025.csv
    • Academic Engstrom, Ho, Sharkey, Cuellar, Government by Algorithm: Artificial Intelligence in Federal Administrative Agencies (Administrative Conference of the United States, Stanford Law School, NYU School of Law, 2020) https://www.law.stanford.edu/wp-content/uploads/2020/02/ACUS-AI-Report.pdf

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Caseworker documentation & copilots domain page.

Levers available here and the patterns behind them

Documented case histories