ParamergeParamerge

Evidence · The claim ledger

Documented deployment cases63

Every cited claim this site makes in this evidence area, with the sources that ground it. Source keys link back to the full reference lists on the Evidence Registry.

EmpiricalMichigan's MiDAS system auto-adjudicated unemployment-insurance fraud with an extremely high error rate among automated …

Michigan's MiDAS system auto-adjudicated unemployment-insurance fraud with an extremely high error rate among automated determinations, wrongly accusing tens of thousands of people; litigation and court action forced review and compensation.

Sources: michiganag2022, ieeespectruma, aiincidentdatabase, benefitstechadvocacyhubb

Appears on: /domains/cases/michigan-midas

EmpiricalThe Royal Commission into the Robodebt Scheme documented hundreds of thousands of wrongful debts raised by an unlawful i…

The Royal Commission into the Robodebt Scheme documented hundreds of thousands of wrongful debts raised by an unlawful income-averaging method, with the onus placed on recipients to disprove automated assessments.

Sources: royalcommissionintotherobode2023a, royalcommissionintotherobode2023b, prygodiczvcommonwealthofaust2021, lawsocietyjournal, royalcommissionintotherobode

Appears on: /domains/cases/australia-robodebt

EmpiricalThe Robodebt Scheme's end came from outside the deploying institution: the Federal Court approved a class-action settlem…

The Robodebt Scheme's end came from outside the deploying institution: the Federal Court approved a class-action settlement covering roughly 430,000 debts for more than $720M (Prygodicz v Commonwealth (No 2) [2021] FCA 634), around 470,000 wrongful debts were to be repaid, and the Royal Commission (2023) found the scheme unlawful.

Sources: prygodiczvcommonwealthofaust2021, royalcommissionintotherobode2023b, royalcommissionintotherobode2023a, lawsocietyjournal, royalcommissionintotherobode

Appears on: /pan-lab, /domains/cases/australia-robodebt

EmpiricalMiDAS's no-review adjudication mode ran from 2013 to 2015, issuing 60,000+ determinations and wrongly accusing roughly 4…

MiDAS's no-review adjudication mode ran from 2013 to 2015, issuing 60,000+ determinations and wrongly accusing roughly 40,000 people; litigation ended it — the Zynda settlement forced reinstatement of human review of fraud determinations (2017), and the Bauserman class action over false fraud accusations settled for $20M (2022).

Sources: michiganag2022, aiincidentdatabase, benefitstechadvocacyhubb, ieeespectruma

Appears on: /pan-lab, /domains/cases/michigan-midas

EmpiricalIndiana's privatized eligibility modernization produced over a million denials in its early years — many procedural rath…

Indiana's privatized eligibility modernization produced over a million denials in its early years — many procedural rather than substantive — before the state canceled the contract and litigated with its vendor.

Sources: eubanks2018c, governmenttechnologyb, ieeespectrumb

Appears on: /domains/cases/indiana-ibm-eligibility

EmpiricalIndependent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic in…

Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.

Sources: rekenkamerrotterdam2021, lighthousereports2023c, wiredlighthousereports2023, followthemoney, racismandtechnologycenter2023

Appears on: /domains/cases/rotterdam-welfare-fraud, /pan-lab

EmpiricalA large share of Arkansas home-care recipients had care hours cut when algorithmic assessment replaced nurse judgment, a…

A large share of Arkansas home-care recipients had care hours cut when algorithmic assessment replaced nurse judgment, and courts found due-process violations centered on the inability to understand or contest determinations.

Sources: arkansasdepartmentofhumanser2017, elderv2022, calo2021, universityofmichiganihpi, benefitstechadvocacyhuba, centerfordemocracytechnology, aiaaic

Appears on: /domains/cases/arkansas-archoices

EmpiricalEvaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations r…

Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.

Sources: goldhaberfiebertprince2019, centreforsocialdataanalytics2019a, rittenhouse

Appears on: /domains/cases/allegheny-afst, /pan-lab

EmpiricalIllinois's Rapid Safety Feedback flagged thousands of children at 90-percent-or-higher risk of serious harm — beyond any…

Illinois's Rapid Safety Feedback flagged thousands of children at 90-percent-or-higher risk of serious harm — beyond any caseload's capacity to act — while children who died in known-to-system cases had not been flagged; the agency ended its use in 2017.

Sources: chicagotribune2017, theimprint2017, governmenttechnologya

Appears on: /domains/cases/illinois-rapid-safety-feedback

EmpiricalOregon's child-welfare agency dropped its AFST-derived Safety at Screening tool in 2022, citing equity concerns amid nat…

Oregon's child-welfare agency dropped its AFST-derived Safety at Screening tool in 2022, citing equity concerns amid national scrutiny of racial disparity in child-welfare algorithms.

Sources: nprap2022, willametteweek2022

Appears on: /domains/cases/oregon-safety-at-screening

EmpiricalDocumented benefit-automation failures replicated determinations into downstream systems with no independent reconciliat…

Documented benefit-automation failures replicated determinations into downstream systems with no independent reconciliation against the source records — Michigan MiDAS actioned replicated flags and Robodebt reversed the onus onto recipients.

Sources: michiganag2022, ieeespectruma, royalcommissionintotherobode2023b, lawsocietyjournal

Appears on: /pan-lab, /practice/reconcile-copied-records, /practice/record-reconciler

EmpiricalDocumented risk-scoring deployments computed scores from multi-agency administrative records originally collected for ot…

Documented risk-scoring deployments computed scores from multi-agency administrative records originally collected for other purposes, which is the data-protection critique recorded in independent reviews of these systems.

Sources: goldhaberfiebertprince2019, eubanks2018c, lighthousereports2023c

Appears on: /pan-lab

EmpiricalDocumented enforcement systems actioned replicated flags automatically — garnishment and penalties applied before any hu…

Documented enforcement systems actioned replicated flags automatically — garnishment and penalties applied before any human review step in the recorded MiDAS deployment.

Sources: michiganag2022, benefitstechadvocacyhubb

Appears on: /pan-lab, /practice/reconcile-copied-records, /practice/record-reconciler

EmpiricalThe Robodebt Royal Commission documented debts raised from income-averaged derived inputs with the onus placed on recipi…

The Robodebt Royal Commission documented debts raised from income-averaged derived inputs with the onus placed on recipients to disprove the automated assessments.

Sources: royalcommissionintotherobode2023b, lawsocietyjournal

Appears on: /pan-lab

EmpiricalRanked risk lists steered which cases were investigated in documented deployments; the anchoring direction is documented…

Ranked risk lists steered which cases were investigated in documented deployments; the anchoring direction is documented while its magnitude is not published.

Sources: lighthousereports2023c, wiredlighthousereports2023, amnestyinternational2021b, goldhaberfiebertprince2019

Appears on: /pan-lab

EmpiricalIn the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations erred at abou…

In the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations erred at about 85% without human review versus 44% with it.

Sources: michiganag2022, ieeespectruma, aiincidentdatabase, benefitstechadvocacyhubb

Appears on: /pan-lab, /domains/cases/michigan-midas, /practice/reconcile-copied-records

EmpiricalIn the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-…

In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.

Sources: rittenhouse, goldhaberfiebertprince2019, centreforsocialdataanalytics2019a, stapletonetal2022

Appears on: /pan-lab, /domains/cases/allegheny-afst

EmpiricalA peer-reviewed 2024 evaluation of the Allegheny Housing Assessment found that although the tool was substantially more …

A peer-reviewed 2024 evaluation of the Allegheny Housing Assessment found that although the tool was substantially more accurate than the VI-SPDAT survey it replaced and produced similar risk-score distributions across race, it did not reduce the racial disparity in service rates: white single adults were served at about 23.3% versus 19.5% for Black clients.

Sources: cheng2024, alleghenycountydepartmentofh2026a

Appears on: /domains/cases/allegheny-housing-assessment, /pan-lab

EmpiricalAfter a 2025 update to the Allegheny Housing Assessment added a fourth outcome predicting future homelessness, the male …

After a 2025 update to the Allegheny Housing Assessment added a fourth outcome predicting future homelessness, the male share of assigned housing rose from 62% to 76% (and the female share fell from 34% to 24%), reflecting a higher measured one-year homelessness risk among men — an example of an outcome-selection choice reshaping who receives scarce housing.

Sources: alleghenycountydepartmentofh2026b

Appears on: /domains/cases/allegheny-housing-assessment, /pan-lab

EmpiricalThe VI-SPDAT was the dominant U.S. homelessness triage assessment for roughly a decade, adopted in at least 39 states an…

The VI-SPDAT was the dominant U.S. homelessness triage assessment for roughly a decade, adopted in at least 39 states and the District of Columbia by 2015, before its own creators announced its phase-out in December 2020 on equity grounds; a 2019 commissioned racial-equity evaluation across four Continuums of Care found race predicted 11 of 16 subscales and that people of color received statistically significantly lower prioritization scores.

Sources: nationalalliancetoendhomeles2022, orgcodeconsultingiaindejong2020, cinnovationswilkey2019

Appears on: /domains/cases/vi-spdat, /pan-lab

EmpiricalThe VI-SPDAT showed poor test-retest reliability, with most participants scoring higher on re-administration, and poor i…

The VI-SPDAT showed poor test-retest reliability, with most participants scoring higher on re-administration, and poor inter-rater reliability, with scores varying by interviewer and site; its predictive validity for housing outcomes was mixed across studies, positive for the youth version, null for single adults in one study, and positive in another community sample.

Sources: bitfocus2021, nationalalliancetoendhomeles2022, shinnandrichard2022

Appears on: /domains/cases/vi-spdat, /pan-lab

EmpiricalThe U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Veterans Heal…

The U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Veterans Health Administration since 2017, scoring about 6.28 million patients and flagging the top 0.1% at each facility (roughly 6,300 to 6,700 veterans a month, more than 130,000 since 2017); an independent re-analysis of 2018 data found the top-0.1% flag has a positive predictive value near 0.05% and a false-negative rate of about 98% for death by suicide, and a 2024 investigation reported that the model treated being a white man as a stronger risk signal than factors specific to women and excluded military sexual trauma and intimate-partner violence from its variables, a characterization VA has contested by framing the excluded factors as less predictive.

Sources: harris2025, glantz2024, graham2025, u2022b

Appears on: /domains/cases/reach-vet, /pan-lab

EmpiricalTwo Veterans Health Administration evaluations of REACH VET found the program associated with improved proximal outcomes…

Two Veterans Health Administration evaluations of REACH VET found the program associated with improved proximal outcomes — more completed outpatient appointments, more new safety plans, and fewer documented suicide attempts — but not with reduced death by suicide: a 2021 triple-differences study of 173,313 veterans across 141 facilities found no association with suicide or all-cause mortality, and a 2025 follow-up of 266,246 observations replicated the null with all confidence intervals crossing one; both are observational rather than randomized studies.

Sources: mccarthy2021, dent2025c

Appears on: /domains/cases/reach-vet, /pan-lab

EmpiricalKaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electronic healt…

Kaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electronic health record of a large virtual mental-health program that handles more than 5,000 intake visits a month; the model is scored in near-real-time (about a 30-minute delay after an encounter trigger) and, at pre-set thresholds, flags high-risk patients to the intake clinician, routing them into the same suicide-risk-assessment and outreach workflow that a positive self-report screen (the PHQ-9 and Columbia-Suicide Severity Rating Scale) triggers, so the machine flag and the self-report alert are effectively OR-merged. In a study of 1,623,232 intake appointments (2012 to 2022, base rate 0.17 percent) the model reached an area under the ROC curve of 0.77 and its top risk decile captured 48.8 percent of appointments later followed by an attempt, but with a positive predictive value of about 0.8 percent.

Sources: hsin2026, hsin2025, papini2024

Appears on: /domains/cases/kaiser-epic-suicide-risk, /pan-lab

EmpiricalBecause the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake is very lo…

Because the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake is very low (0.17 percent) and the positive predictive value in the top risk decile is about 0.8 percent, the large majority of flagged patients will not attempt suicide in the window, so adding the machine-learning flag as a redundant sensor OR-merged onto the existing self-report screen imports a substantial false-positive and clinician-workload burden at scale — a caution the implementation team itself raised. The implementation reports are feasibility- and design-focused and present no evaluation showing the deployment reduced suicide attempts.

Sources: papini2024, hsin2026, hsin2025

Appears on: /domains/cases/kaiser-epic-suicide-risk, /pan-lab

EmpiricalCrisis Text Line, a national nonprofit crisis service, built an in-house machine-learning severity-triage model that reo…

Crisis Text Line, a national nonprofit crisis service, built an in-house machine-learning severity-triage model that reorders which texters volunteer counselors reach first; from about 2017 to 2020 the same anonymized crisis-conversation corpus was routed to Loris.ai, a for-profit spinoff CTL held an ownership stake in — reported by Politico-derived reporting at roughly 53% — which used it to train commercial customer-service software. After a January 28, 2022 Politico exposé, CTL ended the arrangement within three days and requested that the data be deleted; an FCC commissioner referred the matter to the FTC in March 2022, and no public FTC enforcement action is documented. CTL states the shared data was anonymized and never sold as personally identifiable information, and the exact number of records shared has not been made public.

Sources: crisistextlinewikipedia2026, crisistextline2022, reierson2022, bentoninstituteforbroadbanda2022

Appears on: /domains/cases/crisis-text-line-loris, /pan-lab

EmpiricalCrisis Text Line obtained consent for its data collection through an automated reply directing texters to a lengthy Term…

Crisis Text Line obtained consent for its data collection through an automated reply directing texters to a lengthy Terms of Service — described as a roughly 50-page or 4,000-plus-word document — accepted at the moment of acute crisis by users who include many minors; critics including a former board chair, who voted for the data-sharing arrangement and later said she would not have "knowing what I know now," and a terminated volunteer argued that a Terms of Service is not meaningful informed consent for people in crisis. CTL says texters must consent to its privacy policy to use the service and can request deletion by texting the word DELETE, and that since 2023 its in-house research has been overseen by an Institutional Review Board.

Sources: markkulacenterforappliedethi2022, eysenbach2025, reierson2022, trujillo2025

Appears on: /domains/cases/crisis-text-line-loris, /pan-lab

EmpiricalNarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Prescription …

NarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Prescription Drug Monitoring Programs and returns three Narx Scores plus a composite Overdose Risk Score (each 000-999) into the electronic health record, the PDMP portal, or pharmacy software, often in the patient header alongside vitals and allergies; adoption figures vary by what is counted (more than 40 states and territories run their PDMPs on Bamboo technology and five of the top six pharmacy chains use NarxCare, while the scoring module itself is switched on in more than 20 states). The vendor states the scores are intended to aid, not replace, clinical judgment and should never be sole justification for providing or refusing medication, but clinician and patient-advocacy sources document de facto determinative use — denials, forced tapers, and pharmacy refusals — driven by automation bias and fear of regulatory and criminal liability; patients cannot see, challenge, or correct their scores, the algorithm is proprietary and has not been independently validated for clinical care, and the FDA has not regulated it as a Software-as-a-Medical-Device, so contestation has instead run through FDA citizen petitions (one rejected on procedural grounds in 2023 and a second, docket FDA-2025-P-0701, pending since 2025 with more than 1,000 public comments).

Sources: bamboohealth2023, wang2026, millerandwhitehead2023, buonora2023, oliva2022, painnewsnetwork2023, medscape2025a, medscape2025b

Appears on: /domains/cases/narxcare, /pan-lab

EmpiricalOn its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of about 75% (…

On its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of about 75% (self-reported, never independently reproduced), and its own external-validation set from 2017-2023 showed precision falling to about 52%, which the vendor attributed to rising illicit fentanyl (untracked by prescription-monitoring programs) and wider use of opioid-use-disorder treatment medication. A 2026 npj Digital Medicine study that reconstructed the model on California's CURES prescription database (about 17.9 million observations) and on commercial claims data obtained a precision of only 0.01 to 0.32 across several model architectures; because overdose-death labels were unavailable to the independent researchers, that reconstruction was trained on proxy outcomes rather than the score's actual overdose-death target, so it is best read as evidence that proprietary opacity prevents anyone outside the vendor from assessing the deployed model's accuracy, fairness, or safety, rather than as a strict like-for-like refutation of the vendor's figure.

Sources: bamboohealth2023, wang2026

Appears on: /domains/cases/narxcare, /pan-lab

EmpiricalLimbic Access, a Class IIa UKCA-certified self-referral and triage chatbot for NHS Talking Therapies, is deployed across…

Limbic Access, a Class IIa UKCA-certified self-referral and triage chatbot for NHS Talking Therapies, is deployed across a large and growing share of the service (its maker's chief executive claimed about 63% of the NHS in April 2026). Two peer-reviewed observational studies report large operational gains — a study of 129,400 self-referrers across 28 services found referrals rose 15% in chatbot services versus 6% in control services, and a study of 64,862 patients reported clinical-assessment time cut from 54.4 to 41.6 minutes and recovery rates of 58% versus 27.4% — but both studies are non-randomized and were authored by people employed by or holding shares in the tool's maker (all six authors of the access study and seven of the eight authors of the efficiency study), and the efficiency study's own authors caution that the recovery difference is subject to unmeasured confounding from self-selection. No randomized or independent third-party effect estimate has been published.

Sources: habicht2024, rollwage2023, chatterjee2026

Appears on: /domains/cases/limbic-access-nhs, /pan-lab

EmpiricalIn the peer-reviewed study of 129,400 self-referrers across 28 NHS Talking Therapies services, self-referrals rose more …

In the peer-reviewed study of 129,400 self-referrers across 28 NHS Talking Therapies services, self-referrals rose more where the chatbot was in use than in control services (15% versus 6%), with the largest increases among under-served groups — reported at about +179% for nonbinary people, +40% for Black and +39% for Asian self-referrers. This is an observational multi-site association, not a randomized causal effect.

Sources: habicht2024, heikkila2024

Appears on: /domains/cases/limbic-access-nhs, /pan-lab

EmpiricalWoebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people over its l…

Woebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people over its lifetime, was deliberately retired by its maker on a pre-announced schedule: the app was taken down on June 30, 2025, with a transcript-request window (deadline July 15, 2025) and all account data anonymized as of July 31, 2025, removing personally identifying information rather than silently abandoning the service. The founder and chief executive attributed the shutdown to the cost of meeting FDA marketing-authorization requirements and to a regulatory-pathway gap, framing the exit as economic and regulatory rather than a clinical failure - a self-reported account, not an independently audited finding. The roughly 1.5 million figure is a cumulative lifetime number reported in press coverage, not an audited point-in-time active-user count.

Sources: aguilar2025, woebothealth2025, hlth2025

Appears on: /domains/cases/woebot-shutdown, /pan-lab

EmpiricalThe foundational study in Woebot's peer-reviewed efficacy record is an early-stage, vendor-authored trial: a 2017 random…

The foundational study in Woebot's peer-reviewed efficacy record is an early-stage, vendor-authored trial: a 2017 randomized controlled trial in JMIR Mental Health (n=70, ages 18 to 28, two weeks, unblinded, information-only control) reported a moderate between-groups reduction in PHQ-9 depression symptoms (about d = 0.44). That is an efficacy signal, not regulatory validation; the study authors were affiliated with the tool's maker, and no independent, arm's-length evaluation of the consumer app is documented (the broader published evidence base is not assembled here). A separate, investigational, prescription-only variant (WB001) received an FDA Breakthrough Device Designation in May 2021 - an expedited-review status, not marketing authorization - and entered a pivotal Software as a Medical Device trial with the first patient enrolled in January 2023, but never received FDA marketing authorization; it must not be conflated with the consumer app.

Sources: fitzpatrick2017, woebothealthbusinesswire2021b, woebothealthbusinesswire2023

Appears on: /domains/cases/woebot-shutdown, /pan-lab

EmpiricalDWP's own fairness assessment (covering 1 April 2024 to 31 March 2025) of its live Universal Credit Advances fraud-risk …

DWP's own fairness assessment (covering 1 April 2024 to 31 March 2025) of its live Universal Credit Advances fraud-risk model reports statistically significant referral disparities and an accuracy inversion: relative to a 35-44 comparator, claimants aged 55-65 were about 2.80 times as likely to be referred for review and non-UK nationals about 2.27 times as likely, while for older claimants those referrals were less likely to be correct (relative correct-referral likelihoods of about 0.58 at 55-65 and 0.23 at 66-plus, the latter resting on a small sub-sample DWP flags to treat with caution). The disparities were first disclosed under freedom-of-information law and reported in December 2024, and DWP has committed to retrain the model. The figures are DWP-reported relative ratios, not independently audited absolute error rates.

Sources: departmentforworkandpensions2025c, theguardian2024

Appears on: /domains/cases/uk-dwp-fraud-ml, /pan-lab

EmpiricalDWP states that a human caseworker always makes the final decision on a referred Universal Credit advance with no automa…

DWP states that a human caseworker always makes the final decision on a referred Universal Credit advance with no automated decision-making, and is deliberately not shown the risk score or told the referral came from the model; DWP describes the model as around three times more effective than a randomised control at identifying fraud risk and judges continued operation reasonable and proportionate while committing to retrain it. The Public Law Project counters that only age was fully assessed among protected characteristics and that the assessment relied on safeguards preventing downstream harm rather than showing the model to be non-discriminatory. The wider counter-fraud programme is meanwhile expanding into bank-data eligibility verification under the Public Authorities (Fraud, Error and Recovery) Act 2025, a distinct system not yet in force.

Sources: centraldigitalanddataoffice2025, departmentforworkandpensions2025c, publiclawproject2025, departmentforworkandpensions2025b

Appears on: /domains/cases/uk-dwp-fraud-ml, /pan-lab

EmpiricalDuring the pandemic unemployment surge, a private facial-recognition identity check operated as a de facto eligibility g…

During the pandemic unemployment surge, a private facial-recognition identity check operated as a de facto eligibility gate for unemployment benefits in at least 25 U.S. state workforce agencies, with a live 'trusted referee' interview queue that House investigators documented averaging nearly 10 hours in North Dakota and over 4 hours in 14 of 21 states, versus about 6 minutes in New Jersey where an in-person option existed. Oregon's own one-month study (n=10,656 routed) recorded verification-completion differences by group -- for example 41.59% for African American and 34.48% for Spanish-language claimants versus 53.44% for White claimants -- but stated the study showed differences in completion and did not show causation, so these are a friction proxy, not a measured wrongful-denial rate. The U.S. Department of Labor does not collect or report the number of workers blocked for inability to verify identity, and where verification precedes filing those workers are not counted as denied claims at all, so the scale of any wrongful lockout is undocumented.

Sources: ushousecommitteeonoversighta2022a, stateoforegonemploymentdepar2022, nationalemploymentlawproject2023, usdepartmentoflabor2023c

Appears on: /domains/cases/us-idme-unemployment, /pan-lab

EmpiricalA U.S. Department of Labor Inspector General audit (March 31, 2023) found that among 24 state workforce agencies using a…

A U.S. Department of Labor Inspector General audit (March 31, 2023) found that among 24 state workforce agencies using a facial-recognition identity contractor, 18 of 24 (75%) contracts did not specify one-to-one versus one-to-many matching, 15 of 24 (63%) did not address data storage, and 13 of 24 (54%) did not address destruction of the collected biometric data, while 22 of 24 (92%) agencies reported the technology reduced improper payments -- the operator-side benefit that sustained adoption even as the wrongful-lockout cost went unmeasured. The vendor initially represented it used only one-to-one matching and later acknowledged one-to-many matching against a database; after bipartisan backlash the IRS and Treasury dropped the mandatory facial-recognition requirement in February 2022 and the vendor made it optional across agencies, though the service remained in use for unemployment identity verification in a large share of states, and a 2026 IRS proposal would allow it to retain taxpayer biometric data up to 36 months after account deletion. The reported improper-payment reductions are agency self-reports, not independently audited.

Sources: usdepartmentoflabor2023c, americancivillibertiesunionj2022, electronicfrontierfoundation2022, biometricupdate2026

Appears on: /domains/cases/us-idme-unemployment, /pan-lab

EmpiricalOn August 30, 2023 CMS notified states that their automated Medicaid ex parte renewal systems were evaluating eligibilit…

On August 30, 2023 CMS notified states that their automated Medicaid ex parte renewal systems were evaluating eligibility at the household or family level rather than the federally required individual level, so when any one household member could not be auto-renewed the whole household was dropped procedurally if a returned form was not received. CMS found 30 states had the defect and, on September 21, 2023, announced that nearly 500,000 children and other individuals who had been improperly disenrolled would regain coverage, requiring the affected states to pause procedural disenrollments, reinstate coverage, and reprogram to individual-level renewal. The ~500,000 figure is an aggregate of state-reported estimates compiled by CMS, not an independently audited count; children were disproportionately affected because their income-eligibility thresholds are higher than adults', and an HHS ASPE analysis (cited via Georgetown CCF) projected roughly 74% of disenrolled children would still be eligible, a projection rather than a post-hoc audit.

Sources: centersformedicareandmedicai2023a, centersformedicareandmedicai2023b, georgetownuniversitycenterfo2023a, healthcarediveemilyolsen2023

Appears on: /domains/cases/us-medicaid-unwinding-autorenewal

EmpiricalThe Medicaid unwinding was governed by a federal monitor-and-respond loop: Section 5131 of the Consolidated Appropriatio…

The Medicaid unwinding was governed by a federal monitor-and-respond loop: Section 5131 of the Consolidated Appropriations Act, 2023 (SSA section 1902(tt)), codified in a December 6, 2023 interim final rule, gave CMS mandatory monthly state reporting plus, for noncompliance, a Federal Medical Assistance Percentage reduction of 0.25% per quarter (capped at 1%), civil monetary penalties up to $100,000 per day for reporting failure, corrective action plans, and authority to order suspension of procedural disenrollments. Against that instrumentation the overall churn was large: KFF recorded about 25.2 million people disenrolled as of September 12, 2024 with 69% of disenrollments for procedural rather than eligibility reasons, while a June 24, 2025 GAO audit independently found about 27 million disenrolled in the first 18 months, roughly one-third of those continuously enrolled. The KFF and GAO totals differ because they cover different windows and use different data and methods, not because they conflict; the enforcement penalty details are drawn from the interim final rule and a legal-analysis summary.

Sources: federalregister2023, morganlewisandbockiusllp2023, kff2024, usgovernmentaccountabilityof2025

Appears on: /domains/cases/us-medicaid-unwinding-autorenewal

EmpiricalA 2024 Tribunal de Contas da Uniao (TCU) plenary audit found that INSS benefit denials were nonconforming above the maxi…

A 2024 Tribunal de Contas da Uniao (TCU) plenary audit found that INSS benefit denials were nonconforming above the maximum acceptable limit in both channels it sampled: 10.94% of automatically analyzed denials (January to May 2024) and 13.20% of manually analyzed denials (2023 sample), in Acordao 634/2025-Plenario (process TC 008.309/2024-8, session 26 March 2025). Nonconformity ('desconformidade') is a TCU audit-analysis category that includes wrongful denials but is not identical to a court-confirmed wrong-denial rate, so these are not a hard error rate; the absolute counts reported in coverage (about 920,000 automatic denials in the audited window, about 100,000 estimated wrongful, and 250,000 to 290,000 estimated unjustified manual denials) are journalistic extrapolations from the TCU percentages, not officially published counts. Neither channel uses a machine-learning or predictive risk score; 'automatic' means rules-based administrative processing and documentary-conformity analysis.

Sources: tribunaldecontasdauniao2025, infomoney2025, consultorjuridico2025

Appears on: /domains/cases/brazil-inss-automation

EmpiricalThe TCU root-cause finding was that INSS measures server productivity by the number of processes analyzed rather than th…

The TCU root-cause finding was that INSS measures server productivity by the number of processes analyzed rather than the quality of the decision's justification, creating an incentive to choose denial as the fastest disposition, with no incentive for correct motivation of the denial and no effective communication with the insured. The correction channel is slow and external: the CNJ recorded 5,109,076 pending previdenciario lawsuits as of 31 October 2024, and CNJ 'Justica em Numeros' data put the average pending-case duration at about 746 days with a conciliation rate near 24.84%, so a fast automated or manual denial is reversed only after a roughly two-year judicial wait. The 5.1-million-case backlog and the 746-day duration are CNJ caseload figures and cannot be mechanically attributed to automated denials specifically, because the public data do not link an individual court reversal to the channel that produced the denial.

Sources: tribunaldecontasdauniao2025, consultorjuridico2024, conselhonacionaldejustica2024

Appears on: /domains/cases/brazil-inss-automation

EmpiricalThe Commonwealth Ombudsman's first report, Automation in the Targeted Compliance Framework (published 6 August 2025), fo…

The Commonwealth Ombudsman's first report, Automation in the Targeted Compliance Framework (published 6 August 2025), found that the Department of Employment and Workplace Relations and Services Australia acted contrary to the law and unlawfully cancelled the payments of 1,009 jobseekers under the predominantly automated Targeted Compliance Framework, with a further 45 auto-cancelled after a pause was ordered (the first-cohort figure is variously reported as 'more than 900', 964, or 1,009; 1,009 is the most precise and most widely cited). The unlawfulness was an omission: the April 2022 SPROM Act required a discretionary reasonable-excuse consideration before a cancellation and required a mandated automated-decision safeguard, the Digital Protection Framework, neither of which was implemented, so cancellations executed without the check the law required. The defect operated from April 2022, was detected in September 2023 by external legal advisors, and cancellations were not paused until July 2024 — a roughly ten-month gap the Ombudsman called not acceptable; both agencies accepted all seven recommendations. A commissioned Deloitte assurance review separately found the IT system increasingly unstable, with five IT errors dating to 2018, and could not assure the integrity, effectiveness, or appropriateness of decisions.

Sources: commonwealthombudsman2025a, informationageaustraliancomp2025, itnews2025, theexamineraustralianassocia2025, departmentofemploymentandwor2025a

Appears on: /domains/cases/australia-workforce-tcf

EmpiricalThe Targeted Compliance Framework operates at very large scale: advocacy and analysis of departmental data describe roug…

The Targeted Compliance Framework operates at very large scale: advocacy and analysis of departmental data describe roughly 2.5 million payment-suspension notices a year to about a million people, with 200,000 to 240,000 people facing suspension threats each quarter (these system-scale figures are directionally consistent across sources but exact denominators and periods vary). The Ombudsman's second report, Fairness in the Targeted Compliance Framework (published 9 December 2025), found that automatic Penalty-Zone suspensions undermine a jobseeker's ability to challenge penalties, that the department's assessment of provider performance lacks transparency, and that a high rate of provider decisions are overturned on review, while the complaints line went more than 140,000 calls unanswered between November 2024 and September 2025. A separate and far larger section 42AM automated-cancellation review is still expanding: the department previously published up to 9,510 unlawful cancellations or reductions, an advocacy estimate put potential exposure near 310,000, and in June 2026 Senate estimates a department official said the number was in the vicinity of that estimate but qualified that 55 to 70 percent may have legitimately lost eligibility, implying roughly 93,000 or potentially 100,000-plus. Those larger figures are estimates of potentially unlawful cases pending case-by-case assessment, not confirmed cancellations.

Sources: theantipovertycentre2026, powertopersuade2025, commonwealthombudsman2025b, sbsnews2025

Appears on: /domains/cases/australia-workforce-tcf

EmpiricalNew York City launched the MyCity Business chatbot in 2023 on Microsoft Azure AI as a public-facing generative-AI advise…

New York City launched the MyCity Business chatbot in 2023 on Microsoft Azure AI as a public-facing generative-AI adviser for business owners. A March 29, 2024 investigation by The Markup with THE CITY and Documented NY found it confidently and repeatably wrong on legal obligations, advising businesses in ways that would break the law, including that employers could take a cut of workers' tips, that landlords need not accept Section 8 vouchers or source-of-income tenants (illegal in New York City), that stores could go cashless against a 2020 city law, and that funeral-price disclosure could be concealed against the federal funeral rule; when ten staffers asked the housing-voucher question they received the same wrong answer, which had changed from an earlier correct one, showing the tool was non-deterministic. The 2024 findings are qualitative, based on specific tested questions rather than a sampled error rate. The city relabeled the tool a beta product with a disclaimer and applied a scope-narrowing patch rather than withdrawing it, kept it online for roughly two years, and shut it down in early 2026 as a budget cut rather than an accuracy fix.

Sources: themarkup2024, themarkupandthecity2024, reutersjonathanallen2024, themarkupcolinlecherandkatie2026

Appears on: /domains/cases/nyc-mycity-chatbot

EmpiricalA December 30, 2025 performance audit of the MyCity system, issued under New York City Comptroller Brad Lander, found th…

A December 30, 2025 performance audit of the MyCity system, issued under New York City Comptroller Brad Lander, found the chatbot 'appears to be unable to provide accurate or consistent information' and reported that the wider MyCity system had cost over 100 million dollars across more than 120 agreements with about 50 vendors, lacked a system development plan, and had not delivered the promised single-form access to city benefits; the Office of Technology and Innovation disagreed with all seven of the audit's recommendations, including one to conduct AI red-teaming. Among the audit's figures, an internal weekly production report reproduced in the audit showed the chatbot did not answer 23 of 48 tested government questions, and of the more than 2,200 questions asked in July and August 2025 the 70 users who left thumbs-up-or-down feedback were 71.4 percent negative (50 of 70), a share the city disputes as roughly 2.25 percent of all responses, with the audit rebutting that denominator. The 100-million-dollar figure is the whole MyCity system, not the chatbot alone.

Sources: officeofthenewyorkcitycomptr2025

Appears on: /domains/cases/nyc-mycity-chatbot

EmpiricalNevada's Department of Employment, Training and Rehabilitation contracted Google to build a generative-AI tool on the Ve…

Nevada's Department of Employment, Training and Rehabilitation contracted Google to build a generative-AI tool on the Vertex AI Studio cloud platform that reads an unemployment-appeal hearing transcript and evidence, retrieves against a corpus of Nevada unemployment law and prior appeals decisions, and drafts a recommended determination (approve, deny, or modify a claim) together with the written decision for a human referee to review and sign. The contract set a 90 percent success requirement self-assessed by state workers on test decisions -- not an independent external audit -- and DETR said it wanted accuracy higher than 90 percent before going live; rollout was repeatedly delayed over less-than-desired accuracy, including the tool citing incorrect Nevada statutes and failing to pull information from all hearing documents, problems officials said were fixed. Reported cost evolved from about 1 million dollars in 2024 to a total of 2.6 million dollars with about 1.1 million spent by early 2026. As of the most recent available reporting (March 2026) the system was in delayed pre-deployment testing on historical appeals and described as launching in coming weeks; it was not independently confirmed to be adjudicating live claimant appeals.

Sources: thenevadaindependent2025, themarkuptoddfeathers2024, thenevadaindependentericneug2024, fordhamintellectualproperty2024

Appears on: /domains/cases/nevada-detr-genai-appeals, /pan-lab

EmpiricalNevada's generative-AI unemployment-appeals tool was justified as a speed measure for a pandemic-era backlog, projecting…

Nevada's generative-AI unemployment-appeals tool was justified as a speed measure for a pandemic-era backlog, projecting a drop in referee determination time from as much as several hours to about five minutes per case, with a mandatory human review DETR said adds an estimated 10 to 30 minutes and a required referee sign-off (Director Christopher Sewell said no AI-drafted written decisions issue without human review). Legal scholars, attorneys who represent claimants, and a former U.S. Department of Labor official warned that backlog and speed pressure could hollow out that review and create incentives to rubber-stamp AI outputs -- one attorney noting the time savings only happens if the review is very cursory, and a legal analysis warning staff might feel pressured to authorize AI decisions with haste. That automation-deference risk is expert-projected, not a measured outcome: no referee override or rejection rate has been published, and claimants are not required to consent to AI processing of their appeal.

Sources: themarkuptoddfeathers2024, fordhamintellectualproperty2024, thenevadaindependent2025

Appears on: /domains/cases/nevada-detr-genai-appeals, /pan-lab

EmpiricalThe Government Digital Service ran a cross-government experiment with Microsoft 365 Copilot from September 30 to Decembe…

The Government Digital Service ran a cross-government experiment with Microsoft 365 Copilot from September 30 to December 31, 2024, with about 20,000 employees across 12 organisations, and published the findings report on June 2, 2025. Participants self-reported saving an average of about 26 minutes per working day (the report extrapolates this to roughly 13 days a year from the median values of six reported time-savings ranges; independent coverage recomputed it to about 4.6 days on a 253-working-day basis), 17% reported no clear savings, adoption held near 80% after peaking at about 83%, and 82% said they would not want to return to working without it. The experiment measured adoption and self-reported time rather than output quality: the report recorded no audited error rate, flagged significant accuracy concern for low-verifiability tasks such as grievance handling and performance evaluations, noted external web data was used without built-in verification, and documented a provenance failure in which the tool struggled to identify which documents generated a response.

Sources: governmentdigitalservicedsit2025, governmentdigitalservice2025b, theregisterthomasclaburn2025

Appears on: /domains/cases/gds-m365-copilot-experiment, /what-ai-can-do

EmpiricalA companion Department for Business and Trade evaluation of Microsoft 365 Copilot (1,000 licences, October to December 2…

A companion Department for Business and Trade evaluation of Microsoft 365 Copilot (1,000 licences, October to December 2024; published August 28, 2025) reported 72% user satisfaction but concluded it did not find robust evidence that time savings were leading to improved productivity; in observed tasks its users completed spreadsheet data analysis more slowly and to worse quality and accuracy than non-users, and produced presentation slides over 7 minutes faster on average but to worse quality and accuracy that then needed correction. In its diary study, 22% of respondents said they had identified hallucinations, 43% detected none, and 11% were unsure, with a further roughly one in five not answering, so the figure reflects user-detected hallucination rather than audited incidence. A Department for Work and Pensions evaluation (3,549 licences; published January 29, 2026) measured 19 minutes a day saved across eight routine tasks against a comparison group (95% confidence interval 17 to 22 minutes), found 85% rating meeting-note accuracy good or very good, reported that users consistently reviewed outputs before use, and concluded the tool is complementary to human expertise and requires consistent human oversight.

Sources: departmentforbusinessandtrad2025b, theregisterpaulkunert2025, departmentforworkandpensions2026

Appears on: /domains/cases/gds-m365-copilot-experiment, /what-ai-can-do

EmpiricalAccording to its Algorithmic Transparency Recording Standard record, published on November 27, 2025, the UK Department f…

According to its Algorithmic Transparency Recording Standard record, published on November 27, 2025, the UK Department for Work and Pensions runs a Whitemail Insights and Vulnerability Scanner that reads roughly 25,000 scanned documents a day (reported as around 22,000 a day at end-2023 and in a March 2024 operator interview). Each document is passed first through the Vulnerability Scanner, a pre-trained open-source transformer doing zero-shot classification, which flags potentially vulnerable customers against eight prescribed themes including suicide and self-harm, domestic violence and abuse, and financial hardship; only documents not flagged as indicating vulnerability are relayed to Whitemail Insights for routing across nine themes. The output to trained staff is an anonymised daily report of flagged customers, and DWP states the tool does not make or influence benefit entitlement decisions. The record names precision, recall and F1-score as its evaluation metrics but discloses no values, and no independent accuracy evaluation has been published.

Sources: departmentforworkandpensions2025b, trendall2025, ukparliamentworkandpensionsc2023, corbridge2024a

Appears on: /domains/cases/dwp-whitemail-scanner

EmpiricalGuardian FOI reporting in January 2025 recorded that benefit claimants are not told the AI reads their correspondence: t…

Guardian FOI reporting in January 2025 recorded that benefit claimants are not told the AI reads their correspondence: the internal data protection impact assessment stated that letter writers do not need to know about their involvement in the initiative, and the tool had been piloted since at least 2023 without appearing on the central government AI transparency register despite a ministerial mandate. The correspondence it processes can include national insurance numbers, health information, bank details, and children's details. Turn2us policy manager Meagan Levin voiced serious concerns, noting that prioritising some cases inevitably deprioritises others, so it is vital to understand how these decisions are made and ensure they are fair. The further reading that a missed flag on the unflagged residual therefore has no complaint channel and surfaces only as downstream harm is an analytical inference from the documented non-notification and shortlist design, not an adjudicated harm.

Sources: booth2025, toth2025, dent2025a

Appears on: /domains/cases/dwp-whitemail-scanner

EmpiricalIn the UK Home Office's own pilot of an AI tool that summarises asylum interview transcripts for decision-makers, 9% of …

In the UK Home Office's own pilot of an AI tool that summarises asylum interview transcripts for decision-makers, 9% of the generated summaries were deemed inaccurate or incomplete and removed by a pre-use filter before any caseworker saw them, and 23% of users reported not being fully confident in the rest; the summaries carried no source references back to the transcript. The official evaluation, published April 29, 2025, measured a 23-minute-per-case time saving (a 32% reduction) for the summariser and about 37 minutes for a companion policy-search tool, and Home Office Calibre quality-assurance reviews found no statistically significant difference in decision quality on small pilot samples. The evaluation recommended addressing the identified limitations before a full rollout, continuous monitoring in early rollout, and a larger-scale evaluation after deployment; the Home Office announced expansion the same day. By January 2026 the policy-search tool had been rolled out to all asylum decision-makers, and per trade-press reporting the summarisation tool entered national rollout in April 2026.

Sources: ukhomeofficegovuk2025, openrightsgroup2026c, governmenttransformation2026

Appears on: /domains/cases/home-office-asylum-summarisation

EmpiricalThe Home Office's asylum interview-summarisation tool inserts a compression step whose measured value is a 23-minute-per…

The Home Office's asylum interview-summarisation tool inserts a compression step whose measured value is a 23-minute-per-case time saving that exists only insofar as the decision-maker does not redo the reading the summary replaced: caseworkers are not required to verify summaries against transcripts, and the pilot summaries carried no source references that would make checking cheap. The correction loop is also severed from the other side. In a May 2026 written parliamentary answer, minister Alex Norris confirmed that asylum claimants are not told about the AI tools used in their cases, so the one party with first-hand knowledge of their own account cannot surface a summary error; this postdates Article 22C of UK GDPR (in force February 5, 2026). As of mid-2026 the rollout had proceeded without a published post-deployment evaluation or continuous-monitoring data and, per Open Rights Group, without a published Data Protection Impact Assessment, Equality Impact Assessment, or Algorithmic Transparency Recording Standard entry, with prompts withheld under a Freedom of Information refusal. A March 16, 2026 commissioned legal opinion argues the use is likely unlawful on procedural-fairness and data-protection grounds; that is a contested legal position, not a court ruling.

Sources: ukhomeofficegovuk2025, resultsense2026, openrightsgroup2026a, openrightsgroup2026b

Appears on: /domains/cases/home-office-asylum-summarisation

EmpiricalIn February 2026 the Superior Court of Los Angeles County, the largest trial court in the United States, began a pilot o…

In February 2026 the Superior Court of Los Angeles County, the largest trial court in the United States, began a pilot of the Learned Hand AI drafting workbench with six civil-division judges and their research attorneys under a contract of about $314,000 running into early 2027, and the Superior Court of Riverside County gave seven civil and probate research attorneys access under a separate $10,000 agreement used for research memos; the tool ingests case filings, synthesizes applicable law, and drafts proposed orders in the individual judge's own writing style. Under California Judicial Council Rule 10.430 (effective September 1, 2025, the first statewide court generative-AI framework in the nation), disclosure is required only when a document consists entirely of generative-AI output, and the rule reaches judicial officers only for tasks outside their adjudicative role, so neither court is obligated to tell litigants when AI assisted with an order or memo in their case; both courts declined to confirm whether litigants whose cases are used in testing are informed.

Sources: mihalovichandjohnson2026, queally2026, judicialcouncilofcalifornia2025

Appears on: /domains/cases/learned-hand-la-courts

EmpiricalIn the Learned Hand pilot the only reported error-correction safeguard is the judge's own review: officers are required …

In the Learned Hand pilot the only reported error-correction safeguard is the judge's own review: officers are required to review and edit each draft before adopting a tentative ruling, and a court spokesman said the assistance does not supplant the judicial officer's independent role. No external audit, query logging, or benchmarking regime was reported (legal analysis coverage drew a contrast with Michigan's approach), and no error, edit, or override rate for the tool has been published. The Los Angeles District Attorney raised an anchoring concern, that an AI-generated draft could greatly influence what the judge's position should be before an independent view forms; this is an attributed critique rather than a measured effect, and the vendor's per-sentence Deep Verify hyperlinking and multiple-verification-passes claims are unverified vendor statements.

Sources: queally2026, howell2026, mihalovichandjohnson2026, learnedhandandsuperiorcourto2026

Appears on: /domains/cases/learned-hand-la-courts

EmpiricalThe UK Government Digital Service ran what it called the government's biggest public test of generative AI to date: acro…

The UK Government Digital Service ran what it called the government's biggest public test of generative AI to date: across two gated public pilots (a late-2024 web pilot of 10,136 users asking 23,838 questions, and an autumn-2025 GOV.UK app pilot of 641 users asking 2,670 questions in four weeks), more than 10,000 people asked GOV.UK Chat about 26,000 questions on tax, benefits and visas. Its first 2023 version was held back in findings published January 18, 2024 because, GDS reported, answers did not reach the highest level of accuracy demanded for a site like GOV.UK, including a few cases of hallucination. GDS reports measured answer accuracy rising from 76 percent (its earliest benchmark) to 90 percent by the autumn 2025 pilot, assessed by subject-matter experts plus automated evaluation, an 88 percent answer rate for in-scope questions after a clarifying-questions feature was added, and that 508 attempts to jailbreak the system across the pilots were all prevented by its guardrails; it soft-launched to all GOV.UK app users on March 26, 2026 and officially launched on May 14, 2026. Nearly every one of these figures is self-reported by GDS, the system's operator, and the accuracy denominators and sampling frames are unpublished.

Sources: governmentdigitalserviceinsi2026, governmentdigitalserviceinsi2024a, governmentdigitalservice2026

Appears on: /domains/cases/govuk-chat

EmpiricalGOV.UK Chat is a retrieval-augmented assistant that, per its Algorithmic Transparency Record published October 7, 2025, …

GOV.UK Chat is a retrieval-augmented assistant that, per its Algorithmic Transparency Record published October 7, 2025, answers only from roughly 700,000 vectorised chunks (36.9 GB) of curated official GOV.UK guidance, is instructed to ignore its training data, rejects questions containing phone numbers, emails or card numbers, links every answer back to its GOV.UK source pages with a reminder to verify, and retains question data encrypted for 12 months; GDS states it does not attempt to provide advice and makes no automated decision. GDS's December 2025 vision post frames a content-dependency loop, stating that GOV.UK Chat can only be as good as the content published on GOV.UK by departmental teams. The record's independent evaluation is a jailbreak (security) assessment conducted with the AI Security Institute, alongside the record's own caveat that it is not possible to guarantee no jailbreaking attempts will succeed; there is no independent audit of the accuracy methodology, and GDS's claim that for government-related questions the tool scores higher than widely-used consumer AI assistants is the operator's own comparison.

Sources: departmentforscience2025b, governmentdigitalserviceinsi2025, civilserviceworldjimdunton2026

Appears on: /domains/cases/govuk-chat

EmpiricalFrida is the chatbot at the front line of the Norwegian Labour and Welfare Administration's (NAV) anonymous contact-cent…

Frida is the chatbot at the front line of the Norwegian Labour and Welfare Administration's (NAV) anonymous contact-center chat channel; NAV states it launched in summer 2018 and, as of 2026, that citizens first meet Frida (open 24 hours a day) and can ask it for a human advisor on weekdays between 9:00 and 15:00, with the channel anonymous and no personal information visible to NAV. During the COVID-19 lockdown NAV reported a roughly 250 percent surge in inquiries; the platform vendor's case study reports the chatbot answered more than 270,000 coronavirus-related inquiries and that about 80 percent of enquiries were resolved without escalating to a human, and NAV's own funded research report records nearly 11,000 inquiries in Frida on some days between March and May 2020 with a week-13-2020 peak equal to the capacity of about 230 human advisors, where the vendor and the peer-reviewed EJIS study state about 220. These pandemic figures originate substantially in the vendor's marketing case study and are reported here as vendor claims with the 220-versus-230 source tension left unresolved; the roughly 80 percent containment is a completion or non-escalation rate, not a measure of answer accuracy.

Sources: boost2020, parmiggiani2021, vassilakopoulou2022a, nav2026

Appears on: /domains/cases/frida-nav-norway

EmpiricalThe best-documented property of NAV's Frida chatbot is its chatbot-to-human handover boundary, which the evidence sugges…

The best-documented property of NAV's Frida chatbot is its chatbot-to-human handover boundary, which the evidence suggests behaves as a governance-controlled dial: NAV's funded three-university Frida@work project reports that about one in five conversations transferred to a live human advisor under free channel choice, and only about 30 percent of dialogues transferred when NAV removed the explicit choice between the chatbot and human chat, a regime-specific figure that must be read against the interface in force. Independent chat-log studies document irrelevant answers, omitted information, and three classes of domain-knowledge failure, with the most critical failures occurring when a misunderstanding goes undetected inside a conversation the chatbot completed; no per-answer accuracy or error rate has been published, and the Frida@work project found context survives the handover imperfectly, with citizens often unsure whether they are talking to a person or a machine. NAV's own 2025 channel-use analysis found that chatbot visibility appears not to change contact-center inquiry volumes and attributes the steady post-2019 decline to a bundle of causes (self-service improvements, new application systems, SMS notifications and changed contact-center practices), so the chatbot is not shown to reduce human workload outside the crisis peak.

Sources: parmiggiani2021, verne2022, simonsen2020, mcvey2025

Appears on: /domains/cases/frida-nav-norway

EmpiricalBurokratt is Estonia's national network of public-sector chatbots operated by the Information System Authority: each par…

Burokratt is Estonia's national network of public-sector chatbots operated by the Information System Authority: each participating institution runs its own assistant, a central classifier routes a citizen's query between them and oversees the handover, and from 2025 a shared knowledge module built from the eesti.ee state portal feeds cross-domain answers. RIA's page lists 20 participating organisations and trade press reports 18 integrated; an independent 2025 ethnography drawing on twelve insider interviews (conducted in late 2023, when the system spanned ten institutions) found it marketed as advanced AI while functioning much like an FAQ list, with use differing considerably by institution and low in some. No published session volumes, escalation-to-human rates, or answer-accuracy figures, and no dedicated algorithmic-oversight body, published evaluation framework, or national-audit report on the network, were located in the public record.

Sources: informationsystemauthorityri2025a, govinsider2025, kaun2025

Appears on: /domains/cases/burokratt-estonia

EmpiricalA 2025 survey-vignette experiment in Estonia reported that citizens' intended use of a government chatbot relates to per…

A 2025 survey-vignette experiment in Estonia reported that citizens' intended use of a government chatbot relates to perceived usefulness and trust in the technology, that privacy concerns relate to service-provision uses but not to information-provision uses, and that trust in government, explainability, and the amount of information provided were not related to intended use.

Sources: alishani2025

Appears on: /domains/cases/burokratt-estonia

EmpiricalAlbert France Services was a sovereign, in-house generative AI assistant built by DINUM with ANCT to help France Service…

Albert France Services was a sovereign, in-house generative AI assistant built by DINUM with ANCT to help France Services counter advisers answer citizens' benefits and procedure questions from a curated base of official documents, presented by the Prime Minister as a sovereign French AI in April 2024 and, in a demonstration before him, giving a wrong answer on identity-card cost. Piloted from an initial panel of about sixty volunteer advisers to roughly eighty advisers across more than forty counters (forty-eight at final count per AFP) in six departments over three iterated versions, it was, per a January 12, 2026 AFP dispatch, formally not going to be generalized 'in its current form,' a decision DINUM announced on January 9, 2026 while stating that the majority of Albert-brand projects are sustained and fully operational. No error rate, usage volume or override count for the tool was ever published; AFP reports DINUM's annual AI budget at about 1.2 million euros since 2024 with Albert France Services a minimal share, a figure distinct from and not directly comparable to the union Solidaires Finances Publiques' separate claim of a roughly 1.3 million euro project cost.

Sources: wekafrafpdispatch2026, acteurspublics2026, solidairesfinancespubliques2026, franceservicesanct2024

Appears on: /domains/cases/albert-france-services

EmpiricalFor Albert France Services no instrumented error-detection channel existed: no error rate, override count or usage figur…

For Albert France Services no instrumented error-detection channel existed: no error rate, override count or usage figure was published during the pilot, and the failures that framed the tool surfaced through the operator side, with several unions documenting recurring malfunctions and wrong answers, an investigative-television broadcast in April 2025 (per Solidaires Finances Publiques) featuring unenthusiastic agent testimony, and advisers reporting answers worse than an ordinary search. According to Solidaires Finances Publiques the project had in fact stopped by September 2025 with no announcement, inferred from Albert no longer appearing among projects presented in a ministerial working group (a union claim). Alongside the January 2026 non-generalization decision DINUM migrated the Albert API model aliases off the 'albert-' branding and removed the web-search functionality, retiring legacy aliases by February 15, 2026, while a successor adviser tool that integrates models from the vendor Mistral AI was in test with about 10,000 public agents through June 2026, gated by a summer-2026 evaluation that must notably establish the cost of a generalization.

Sources: solidairesfinancespubliques2026, nextnextink2026, wekafrafpdispatch2026, acteurspublics2026

Appears on: /domains/cases/albert-france-services