Evidence · The claim ledger
Other233
Every cited claim this site makes in this evidence area, with the sources that ground it. Source keys link back to the full reference lists on the Evidence Registry.
EmpiricalAn independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without human over…
An independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without human override, it would have recommended screening in about 68% of Black children versus 50% of white children (an 18-point gap), while call screeners actually screened in 51% and 43% (a 7-point gap) — the narrower gap came from workers disagreeing with the score about a third of the time.
Sources: stapleton, stapleton2025, hoandburke2022
Appears on: /domains/cases/allegheny-afst, /pan-lab
EmpiricalAn ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of Black ref…
An ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of Black referral-households in the data were affected by at least one permanent 'ever-in' variable drawn from public-benefits data sources, compared with 80% of non-Black households.
Sources: gerchicketal2023
Appears on: /domains/cases/allegheny-afst, /pan-lab
EmpiricalThe U.S. Department of Justice's Civil Rights Division was reported to be scrutinizing the Allegheny Family Screening To…
The U.S. Department of Justice's Civil Rights Division was reported to be scrutinizing the Allegheny Family Screening Tool after civil-rights complaints filed in fall 2022 raised concerns that its use of disability, mental-health, and Supplemental Security Income data may discriminate against parents with disabilities; families are not shown their scores, and no public findings or enforcement have been reported.
Sources: associatedpress2023, hoandburke2023
Appears on: /domains/cases/allegheny-afst
EmpiricalIn Allegheny County's Hello Baby program, the top-tier roughly 5% of newborns by predictive risk score accounted for abo…
In Allegheny County's Hello Baby program, the top-tier roughly 5% of newborns by predictive risk score accounted for about 54% of children later removed from the home by age three, at roughly twenty times the removal risk of other newborns (methodology relative risk 22.24, 95% CI 17.50-28.25); the model reported an AUC of about 0.93 on holdout data.
Sources: centreforsocialdataanalytics2020, vaithianathan2025
Appears on: /domains/cases/allegheny-hello-baby
EmpiricalAn external evaluation of Hello Baby covering birth cohorts 2016-2024, controlling for COVID-19, found the program assoc…
An external evaluation of Hello Baby covering birth cohorts 2016-2024, controlling for COVID-19, found the program associated with fewer first child-maltreatment investigations and first substantiated investigations, but no reduction in out-of-home foster-care placements - the outcome the predictive model was built to estimate.
Sources: lery2025, centreforsocialdataanalytics2020
Appears on: /domains/cases/allegheny-hello-baby
EmpiricalIn a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0.95 for a …
In a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0.95 for a two-year outcome — a child's separation from family or contact with child-protection programs — using 280 administrative variables per child; the deployed operational model's real-world performance was never publicly disclosed.
Sources: derechosdigitalesmatiasvalde2021, derechosdigitalesmatiasvalde2022, centreforsocialdataanalytics2019b
Appears on: /domains/cases/chile-alerta-ninez, /pan-lab
EmpiricalSistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits, without …
Sistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits, without informed consent to the risk ranking or a way to opt out; the model's developers acknowledged it was less able to identify higher-income children at risk, because lower-income families have more contact with the state.
Sources: derechosdigitalesmatiasvalde2021, derechosdigitalesmatiasvalde2022, centerforhumanrightsandgloba2022
Appears on: /domains/cases/chile-alerta-ninez, /pan-lab
EmpiricalThe Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019, scores ea…
The Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019, scores each referral from 1 to 20 for a child's likelihood of out-of-home removal within two years; an independent Cornell-led randomized controlled trial found it sped up screening decisions without significantly changing child outcomes, and a companion study found workers attended mainly to extreme scores while largely disregarding mid-range ones.
Sources: vaithianathanetalcentreforso2019, fitzpatrick2025, eiermann2026
Appears on: /domains/cases/douglas-county-decision-aid, /pan-lab
EmpiricalEckerd's Rapid Safety Feedback spread from Hillsborough County, Florida to child-welfare agencies in several states — pr…
Eckerd's Rapid Safety Feedback spread from Hillsborough County, Florida to child-welfare agencies in several states — promoted on the vendor's own reported gains and highlighted as “innovative” in a 2016 federal commission report — years before an independent 2022 peer-reviewed evaluation found the process did not lower repeat high-severity maltreatment among children identified as high risk (a joint odds ratio of about 1.05).
Sources: eckerdconnects2016, routefiftygovernmentexecutiv2016, parker2022, floridaschildrenfirst2012
Appears on: /domains/cases/eckerd-florida-rsf-origin
EmpiricalGladsaxe's early-detection project (DTO) was a decision-tree model over about 44 risk indicators, meant to score, for ev…
Gladsaxe's early-detection project (DTO) was a decision-tree model over about 44 risk indicators, meant to score, for every child aged 0 to 6 rather than only families already receiving help, the estimated probability that the child was living in vulnerability; per a university-run Danish public-sector AI catalogue it was to be trained on roughly 173,000 notifications the authorities received between April 2016 and December 2017, but only about 117 usable historical cases existed, and it was halted in its development phase in 2019 without ever running on live decisions, after a national media storm and an unrelated data breach that exposed about 20,000 citizens' personal identification numbers.
Sources: offentligaiuniversityrundanind, kennethkristensensamfundsled2022, helenefriisratnerandkasperel2023, katarinafastlappalainen2021, tvkosmopolformerlytvlorry2018
Appears on: /domains/cases/gladsaxe-denmark
EmpiricalHackney paid the analytics firm Xantura £361,400 over four years to run an Early Help Profiling System that flagged fami…
Hackney paid the analytics firm Xantura £361,400 over four years to run an Early Help Profiling System that flagged families for preventive intervention from council data, but scrapped the pilot in 2019 after finding that, despite flagging about 350 families, it surfaced only 7 children previously unknown to the council and the available data was too limited and variable to justify continuing.
Sources: hackneycouncilpayskpoundstod2018, townhalldropspilotprogrammep2019
Appears on: /domains/cases/hackney-early-help
EmpiricalFamilies whose data Hackney's Early Help Profiling System processed were not informed directly: reporting describes fami…
Families whose data Hackney's Early Help Profiling System processed were not informed directly: reporting describes families profiled without their knowledge, given notice only through a general online privacy notice, with no option to opt out recorded in the system's impact assessment and the method withheld as commercially sensitive; the council argued that disclosing the system could prejudice potential interventions.
Sources: townhalldropspilotprogrammep2019, reddenj2020, hackneycouncilpayskpoundstod2018
Appears on: /domains/cases/hackney-early-help
EmpiricalInternal DCFS tracking data released under Illinois public-records law showed the Rapid Safety Feedback tool flagged mor…
Internal DCFS tracking data released under Illinois public-records law showed the Rapid Safety Feedback tool flagged more than 4,100 children at a 90-percent-or-higher probability of death or serious injury within two years, including 369 children under age 9 assigned a 100-percent probability, while children who died in cases already known to the system — among them 17-month-old Semaj Crosby, found dead after at least ten DCFS investigations — were not flagged as top-risk; the roughly $366,000 program was ended in 2017.
Sources: chicagotribune2017, governmenttechnologya
Appears on: /domains/cases/illinois-rapid-safety-feedback
EmpiricalIllinois brought in the Eckerd/MindShare Rapid Safety Feedback program under DCFS director George Sheldon through a no-b…
Illinois brought in the Eckerd/MindShare Rapid Safety Feedback program under DCFS director George Sheldon through a no-bid arrangement the state classified as a grant; a July 2017 joint report by the Illinois Office of Executive Inspector General and the DCFS Inspector General found this classification to be mismanagement because it avoided state bidding-transparency requirements.
Sources: chicagotribune2017, sunshinestatenews2017
Appears on: /domains/cases/illinois-rapid-safety-feedback
EmpiricalBristol's Think Family Database drew on roughly 30 to 35 fused council, police and other datasets covering about 55,000 …
Bristol's Think Family Database drew on roughly 30 to 35 fused council, police and other datasets covering about 55,000 families (some 170,000 residents in 2021 reporting), and its child sexual and criminal exploitation risk models were quietly withdrawn in 2023 as 'not fit for operational use' after an independent evaluation judged the risk-scoring models the weakest element and staff reported victims of exploitation scoring below people involved in burglary; FOI responses indicate no record was kept of why the models were switched off, and auditors could not locate their source code or variable lists.
Sources: seanmorrison2026a, markwildingandmattburgess2026, seanmorrison2026b, bristolcitycouncil2025, jakehurfurtbigbrotherwatch2021
Appears on: /domains/cases/insight-bristol
EmpiricalReporting and FOI responses on Bristol's Think Family Database indicate the exploitation models' source code and variabl…
Reporting and FOI responses on Bristol's Think Family Database indicate the exploitation models' source code and variable lists could not be located when auditors sought them, and that an ethics committee advising the police analytics reportedly did not revisit the analytics after 2017; a 2021 review warned that data gathered through 'legal gateways' meant 'legality is not the same as legitimacy.'
Sources: markwildingandmattburgess2026, seanmorrison2026a, seanmorrison2026b
Appears on: /domains/cases/insight-bristol
EmpiricalIn a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built …
In a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built by SAS — correctly flagged 171 of the highest-risk children but produced 3,829 false positives, a false-positive rate of about 95.6% that DCFS's own public-affairs director confirmed on the record, and the county shelved the tool in 2017 without ever using it on a live case.
Sources: theimprintdanielheimpel2015, childprotectiveservicesdefen2015, nccprrichardwexler2017, witnesslarichardwexler2017
Appears on: /domains/cases/la-county-aura, /pan-lab
EmpiricalThe Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow child-ris…
The Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow child-risk flags (902 of 2,444 over three months across four police regions, rising to 53% in Amsterdam-Amstelland) were system or registration errors or based on irrelevant incidents, and that in none of the four regions was there a well-functioning instrument.
Sources: dspgroepforthewodcabraham2011, dimitritokmetzissargasso2012
Appears on: /domains/cases/netherlands-prokid, /pan-lab
EmpiricalNew Zealand's Ministry of Social Development commissioned a child-maltreatment risk-modelling tool that, on a 2012 devel…
New Zealand's Ministry of Social Development commissioned a child-maltreatment risk-modelling tool that, on a 2012 development sample of 57,986 children and 132 selected variables, reported an area under the ROC curve of 76% and a top risk decile in which 47.8% had a substantiated maltreatment finding by age five; those figures come from development data rather than field performance, the tool was never operationally deployed, and a proposed two-year study that would have scored about 60,000 newborns was halted by the incoming Social Development Minister, who annotated the briefing papers 'Not on my watch! These are children not lab rats.'
Sources: vaithianathan2013, nzherald2015, otagodailytimes2015, mordaunt2026
Appears on: /domains/cases/nz-msd-prm
EmpiricalOregon's Department of Human Services stopped using its Safety at Screening tool at the end of June 2022 and replaced it…
Oregon's Department of Human Services stopped using its Safety at Screening tool at the end of June 2022 and replaced it with a non-algorithmic Structured Decision Making process, telling staff the change was meant to reduce disparities; the move followed Associated Press reporting on racial disparity in the Allegheny tool it was derived from and a racial-bias inquiry from a U.S. senator.
Sources: associatedpress2022, willametteweek2022, hoandburke2022, nprap2022
Appears on: /domains/cases/oregon-safety-at-screening
EmpiricalOregon's 2019 report describes a post-processing fairness correction — group-specific thresholds selected under an 'erro…
Oregon's 2019 report describes a post-processing fairness correction — group-specific thresholds selected under an 'error rate balance' criterion — applied to a dual-outcome risk model built only on the state's own child-welfare administrative records.
Sources: orrai2019, associatedpress2022
Appears on: /domains/cases/oregon-safety-at-screening
EmpiricalNone of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities…
None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.
Sources: claytonandsanders2022, communitycareturner2020a, childhubterredeshommes2020
Appears on: /domains/cases/wwc-uk-ml-pilots, /pan-lab
EmpiricalIn a survey of 129 social workers carried out for the project, only about 26% supported using predictive analytics to id…
In a survey of 129 social workers carried out for the project, only about 26% supported using predictive analytics to identify families for early help and about 34% thought it should not be used at all.
Sources: communitycareturner2020a
Appears on: /domains/cases/wwc-uk-ml-pilots, /pan-lab
EmpiricalIn a Los Angeles County pilot, 335 people who enrolled in the voluntary Homelessness Prevention Unit were reported to be…
In a Los Angeles County pilot, 335 people who enrolled in the voluntary Homelessness Prevention Unit were reported to be 71% less likely than a regression-adjusted comparison group of 1,285 eligible non-enrollees to enter a homeless shelter or have street-outreach contact within 18 months; the California Policy Lab describes this as an association not yet shown to be causal, pending a randomized controlled trial with results expected in 2027.
Sources: blackwell2025, countyoflosangeles2025, uclanewsroom2025
Appears on: /domains/cases/la-homelessness-prevention
EmpiricalThe Homelessness Prevention Unit's own November 2024 equity audit, on a test population of 47,582 individuals eligible t…
The Homelessness Prevention Unit's own November 2024 equity audit, on a test population of 47,582 individuals eligible to be scored, reported false-negative rates ranging from about 56% for Black individuals to roughly 63 to 65% for other groups: the model misses a majority of the people who later become homeless, while performing roughly consistently across race, ethnicity, and gender and identifying Black individuals slightly more strongly.
Sources: californiapolicylab2024, foxsowell2025
Appears on: /domains/cases/la-homelessness-prevention
EmpiricalXantura's OneView integrates more than 15 multi-agency data feeds into a single household view and flags residents as li…
Xantura's OneView integrates more than 15 multi-agency data feeds into a single household view and flags residents as likely to become homeless months ahead. In Maidstone's pilot year it produced 650-plus alerts that a single financial-inclusion officer could contact only about 260 of. Its headline effectiveness figures - a reported 40 percent fall in homelessness, savings and an ROI over 600 percent, and the widely quoted contrast between contacted and uncontacted households - are vendor- and council-reported pre/post numbers from one COVID-affected pilot year; the contact-versus-no-contact contrast reflects capacity-driven selection rather than a randomised comparison, and the independent randomised controlled trial commissioned to test the causal claim was still in progress into 2026.
Sources: crisisuk2023, xantura2023, governmenttransformationmaga2023, ministryofhousing2024, centreforhomelessnessimpact2024
Appears on: /domains/cases/xantura-oneview-housing
EmpiricalOneView's single view of vulnerability is built by integrating sensitive multi-agency records - including offending, hea…
OneView's single view of vulnerability is built by integrating sensitive multi-agency records - including offending, health, benefits and debt data - under a statutory Digital Economy Act 2017 data-sharing agreement with named public-body controllers and processors. An independent ethnography of an early deployment (its fieldwork centered on children's social care and the COVID-19 response) found frontline staff could not see which factors drove the tool's alerts and were not all convinced it was as accurate as described, and a separate NGO investigation characterised the vendor's COVID-era model as operating without residents' knowledge.
Sources: digitaleconomyactregister2023, adalovelaceinstitute2024, bigbrotherwatch2021
Appears on: /domains/cases/xantura-oneview-housing
EmpiricalLondon, Ontario's CHAI is a live, caseworker-facing machine-learning model that flags people in the city's shelter syste…
London, Ontario's CHAI is a live, caseworker-facing machine-learning model that flags people in the city's shelter system as at risk of chronic homelessness (more than 180 shelter days in a year) about six months ahead; it provides intelligence to prevention caseworkers and does not itself make service decisions. Its widely repeated '93 percent accuracy' is a builder-reported, testing-phase figure from 10-fold cross-validation on historical HIFIS records, never independently validated after deployment; the same technical work reports recall of about 0.921 but precision of only about 0.651, implying substantial false positives under a low base rate.
Sources: wray2020, vanberlo2009, lebel2023, govlaunchstories2020
Appears on: /domains/cases/chai-london-ontario
EmpiricalCHAI is consent-based: it draws on de-identified HIFIS records pooled from roughly 20 to 24 London homelessness-support …
CHAI is consent-based: it draws on de-identified HIFIS records pooled from roughly 20 to 24 London homelessness-support organizations and lets individuals opt out of inclusion, and it was built with reference to GDPR principles, Canada's Directive on Automated Decision-Making, and local feature-attribution explanations for caseworkers. Because HIFIS captures people who use public shelters, an independent review and reporting at launch note it can under-represent or miss groups who avoid them - including many women, families, new immigrants, some Indigenous people, and private-shelter users; academic researchers situating the tool raise related fairness and inequality concerns. So the population the model can score is a selected sample of actual need, and the opt-out self-selects it further.
Sources: wray2020, lebel2023, lamberink2020, redden2026
Appears on: /domains/cases/chai-london-ontario
EmpiricalIn a 2025 Los Angeles County pilot evaluated by Nava Labs with academic partners at Cornell University and Georgetown Un…
In a 2025 Los Angeles County pilot evaluated by Nava Labs with academic partners at Cornell University and Georgetown University's Better Government Lab, a generative-AI assistive chatbot for Imagine LA's Benefit Navigator was estimated to improve benefits-navigation answer accuracy by an average of about 40% in a randomized controlled trial of 125 caseworkers answering hypothetical client questions, alongside a fourteen-week field pilot with 61 caseworkers across six organizations; the evaluation was co-authored by the tool builder rather than independently replicated, the accuracy figure is a decision-support contrast on hypothetical questions rather than a live-caseload eligibility audit, and time-savings and administrative-burden effects were reported as promising but inconclusive (published March 2026).
Sources: navapublicbenefitcorporation2026, chen2026
Appears on: /domains/cases/imagine-la-benefit-navigator
EmpiricalThe same evaluation reported that the chatbot's accuracy gains were largest on the most difficult client questions and a…
The same evaluation reported that the chatbot's accuracy gains were largest on the most difficult client questions and among the newest, least-experienced staff (a directional finding, not a quantified breakdown), that about 65% of caseworkers with access used it at an average of about 14 prompts each and a modest, low-positive satisfaction (a Net Promoter Score of 11), that usage tended to decline over time without sustained engagement, and that answers averaged a tenth-to-twelfth-grade reading level against college-level source manuals.
Sources: navapublicbenefitcorporation2026, chen2026
Appears on: /domains/cases/imagine-la-benefit-navigator
EmpiricalIn a single-center randomized trial across three Vanderbilt neurology clinics (August 2022 to February 2023), an EHR sui…
In a single-center randomized trial across three Vanderbilt neurology clinics (August 2022 to February 2023), an EHR suicide-risk model flagged 596 of 7,732 encounters (about 8%) at a 2%-or-higher 30-day-risk threshold; making the identical alert interruptive rather than passive led clinicians to elect a suicide-risk screen in 42% of encounters (121/289) versus 4% (12/307) for a passive chart icon, an adjusted odds ratio of 17.70 (95% CI 6.42–48.79). Screening remained fully advisory: about 58% of interruptive and 96% of passive alerts produced no screening.
Sources: walshetal2025, aitestedforalertingclinician2025, suicidepreventionmorefeasibl2025
Appears on: /domains/cases/vsail-vanderbilt
EmpiricalIn a separate 2021 prospective silent-mode study (115,905 predictions on 77,973 patients, June 2019 to April 2020), the …
In a separate 2021 prospective silent-mode study (115,905 predictions on 77,973 patients, June 2019 to April 2020), the model reported a c-statistic of 0.797 for suicide attempt and 0.836 for ideation center-wide but only 0.544 for attempt in behavioral-health settings, and in the highest-risk quantile the number-needed-to-screen was 271 for attempt and 23 for ideation. In the 2022 to 2023 trial no suicidal ideation or attempts were documented in either arm during 30-day follow-up, and the trial was explicitly not powered for clinical outcomes, so it measured a process outcome (screening) rather than reduced harm.
Sources: walshetal2021, walshetal2025, suicidepreventionmorefeasibl2025
Appears on: /domains/cases/vsail-vanderbilt
EmpiricalBetween roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrongly accused…
Between roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrongly accused an estimated 26,000 or more families of childcare-benefit fraud and demanded full repayment; broader advocacy estimates run higher and count different populations, and by February 2026 about 69,000 people had applied to the recovery scheme and more than 43,000 were formally recognized as affected, each entitled to a minimum of 30,000 euros. A self-learning risk-classification model that scored applications using a Dutch-nationality indicator, a 270,000-person fraud blacklist (the FSV) held without a legal basis, and an all-or-nothing recovery regime were coupled together; the Dutch Data Protection Authority imposed 6.45 million euros in fines (2.75 million for the nationality processing in 2021 and 3.7 million for the FSV blacklist in 2022), a parliamentary inquiry found rule-of-law violations, and the third Rutte cabinet resigned on 15 January 2021.
Sources: wikipedia2026, autoriteitpersoonsgegevens2021, autoriteitpersoonsgegevens2022, amnestyinternational2021b, tweedekamerderstatengeneraal2020, rijksoverheid2026
Appears on: /domains/cases/nl-toeslagenaffaire, /pan-lab
EmpiricalThe scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. Government-…
The scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. Government-commissioned technical reviews (KPMG in 2022 and PwC in 2023) described the tool as a self-learning classifier that routed the highest-scoring of roughly 90,000 benefit applications sent to manual treatment in 2014 to 2019, but judged the Dutch-nationality indicator's standalone predictive weight to have been limited; the model's precision and false-positive rate were never measured or published. The FSV fraud blacklist held frequently inaccurate data that was not corrected when people were cleared, and internal 2016 guidance auto-labelled childcare debts over 3,000 euros as intent or gross negligence, blocking payment arrangements. Out-of-home child placements are a documented but causally contested downstream harm: statistics counted roughly 2,090 children of affected parents placed out of home through mid-2022, while a 2025 judicial study found no child was removed solely because of financial problems.
Sources: kpmg2022, pwc2023, autoriteitpersoonsgegevens2022, statisticsnetherlandscbs2022, rechtspraak2025, wikipedia2026
Appears on: /domains/cases/nl-toeslagenaffaire, /pan-lab
EmpiricalOn 5 February 2020 the District Court of The Hague ruled that the legislation authorising SyRI, the Dutch state's secret…
On 5 February 2020 the District Court of The Hague ruled that the legislation authorising SyRI, the Dutch state's secret cross-database welfare-fraud risk-profiling system, violated Article 8 of the European Convention on Human Rights, and it ordered the system's use stopped; the State did not appeal. The ruling is widely described as one of the first times a court anywhere halted a digital welfare-fraud technology on human-rights grounds. Across its two executed neighbourhood projects SyRI was reported to have produced no confirmed fraud cases, and in one municipality 62 of 113 risk notifications were reported to be false positives.
Sources: districtcourtofthehague2020, vanbekkum2021, unofficeofthehighcommissione2020, algorithmwatch2020a, pontdataprivacyprivacywebnl2019, publicinterestlitigationproj2020
Appears on: /domains/cases/nl-syri
EmpiricalThe District Court of The Hague found that the SyRI framework provided no duty to notify people that their data had been…
The District Court of The Hague found that the SyRI framework provided no duty to notify people that their data had been processed or that a risk report had been filed, so a flagged person generally could not know about, access, or contest the notification; notifications were retained in a register for up to two years. The court held that a risk notification carried significant effect for the person even though it lacked formal legal effect, and it faulted the scheme for a lack of transparency and for breaching data-minimisation and purpose-limitation principles.
Sources: districtcourtofthehague2020, vanbekkum2021
Appears on: /domains/cases/nl-syri
EmpiricalFrance's family-benefits fund (CNAF) computes a monthly benefit-fraud suspicion score, on a 0-to-1 scale, for every bene…
France's family-benefits fund (CNAF) computes a monthly benefit-fraud suspicion score, on a 0-to-1 scale, for every benefit-receiving household — analysing the data of about 32 million people and producing more than 13 million scores each month, close to half of France's population; the highest scores route households into fraud controls, up to the most invasive on-site checks. An analysis by Le Monde and Lighthouse Reports of an extracted production model (a logistic regression of about 33 variables) found that markers of economic vulnerability raised the score: a stable-income family averaged about 0.33, while a person working while receiving the disability allowance (AAH) averaged about 0.66. The model's target was an overpayment (indu) above a threshold, which is frequently unintentional administrative error rather than proven intentional fraud, and the score itself is not disclosed to the person and cannot be appealed directly. CNAF disputed the discrimination framing, describing the tool as a neutral decision-aid that only prioritises which files to check; a coalition that grew to 25 organisations challenged the model before the Conseil d'État, and as of this writing no court had ruled.
Sources: lighthousereports2023b, lighthousereports2023a, laquadraturedunet2023, laquadraturedunet2026a, amnestyinternational2024b, generationnt2026
Appears on: /domains/cases/france-cnaf
EmpiricalIn an internal simulation study by CNAF's own statistics department (DSER), reported in October 2025 by Le Monde and La …
In an internal simulation study by CNAF's own statistics department (DSER), reported in October 2025 by Le Monde and La Quadrature du Net, recipients of the RSA minimum-income benefit were about 13% of beneficiaries but 39 to 41% of the highest-scoring 5%, and single mothers were about 14% of beneficiaries but 37 to 40% of that top bracket; households including a foreign national scored higher on average even after the nationality variable was removed. The full study is not public, and false-positive rates by protected group have not been released. The French ombudsperson (Défenseur des droits) told the Conseil d'État that a presumption of indirect discrimination appeared established because the differential treatment rests on beneficiaries' economic vulnerability; CNAF disputed the characterisation, and no court had ruled.
Sources: laquadraturedunet2026b, generationnt2026, laquadraturedunet2026a
Appears on: /domains/cases/france-cnaf
EmpiricalAnalysing the Swedish Social Insurance Agency (Forsakringskassan) 2017 outcome data, Lighthouse Reports and Svenska Dagb…
Analysing the Swedish Social Insurance Agency (Forsakringskassan) 2017 outcome data, Lighthouse Reports and Svenska Dagbladet reported on 27 November 2024 that the agency's in-house machine-learning risk profile for the temporary parental allowance (VAB) selected women (more than 1.5x), people of a foreign background (about 2.5x), below-median earners (2.97x), and people without a university degree (3.31x) for fraud investigation more often than comparison groups by demographic parity, and wrongly flagged those groups at higher false-positive rates (about 1.7x for women and 2.4x for people of a foreign background); in the agency's paired random-control sample, 20.2 percent of applications contained at least one day incorrectly paid, an unbiased base error rate. These are outcome computations under specific fairness definitions from a single obtained year of data, not confirmed model internals; the agency disputed the framing and did not release the model. The data-protection regulator IMY closed its GDPR supervision on 18 November 2025 for mootness after the agency withdrew the system, and no court or regulator issued a discrimination or GDPR penalty.
Sources: lighthousereports2024, lighthousereportsb, lighthousereportsc, integritetsskyddsmyndigheten2025a, integritetsskyddsmyndigheten2025b
Appears on: /domains/cases/sweden-forsakringskassan
EmpiricalThe Swedish Social Insurance Agency (Forsakringskassan) did not disclose the machine-learning risk profile it used to se…
The Swedish Social Insurance Agency (Forsakringskassan) did not disclose the machine-learning risk profile it used to select temporary-parental-allowance recipients for fraud investigation: its algorithm class, features, and precision were never released, and the agency resisted freedom-of-information disclosure for roughly three years on fraud-prevention grounds. In 2018 the audit inspectorate ISF found the risk-based profiling substantially more accurate than alternative controls while warning that it raised legal-certainty and equal-treatment concerns, and cautioning that an accurate model can still be inequitable when two groups err equally but only one is followed up. Amnesty International reported that a former agency data protection officer warned in 2020 that the operation breached European data-protection rules. The system was decommissioned in 2025 during the regulator's supervision, before any court or regulator ruled on it.
Sources: lighthousereports2024, lighthousereportsb, inspektionenforsocialforsakr2018a, inspektionenforsocialforsakr2018b, amnestyinternational2024d
Appears on: /domains/cases/sweden-forsakringskassan
EmpiricalDenmark's Udbetaling Danmark (UDK), administered by ATP, runs a data-driven welfare-fraud operation that as of 2019 used…
Denmark's Udbetaling Danmark (UDK), administered by ATP, runs a data-driven welfare-fraud operation that as of 2019 used up to about 60 AI and machine-learning models to score benefit recipients into a 'wonderlist' of high-risk people, which a human control team filters into control cases for investigation. In UDK's own 2023 control statistics (three documented models), the 'Model Abroad' foreign-affiliation model sent 511 cases for control but recovered money in only 36 -- about 7%, with roughly nine in ten resulting in no further action -- and UDK confirmed that 54% of the 'Really Single' household-outlier cases its unit opened were in fact legitimate. Those 'revenue' outcomes conflate deliberate fraud with honest error, which UDK does not separate, so they are not pure fraud rates. Amnesty International characterised the system as mass surveillance and prohibited social scoring under the EU AI Act; UDK, ATP and the ministry (STAR) rejected that characterisation, the system was not suspended, and as of this writing no court had ruled.
Sources: amnestyinternationalalgorith2024, amnestyinternational2024a, fortuneeurope2024, bablai2024
Appears on: /domains/cases/denmark-udbetaling
EmpiricalUdbetaling Danmark's 'Joint Data Unit' merges and links the personal data of millions of residents from around ten natio…
Udbetaling Danmark's 'Joint Data Unit' merges and links the personal data of millions of residents from around ten national registers -- civil registration (CPR), buildings and dwellings (BBR), business, income, tax (R75), health, VAT, cash and sickness benefits, education grants and the motor-vehicle register -- alongside a 'Joint Data Unit Abroad' that pulls data from foreign authorities; in 2021 UDK paid about DKK 241 billion to roughly 2.4 million recipients. Amnesty International documents this as mass surveillance and argues the design carries a discrimination risk: 'Model Abroad' scores a relative strength of ties to non-EEA countries with citizenship as a direct input, and 'Really Single' treats statistically atypical households as suspicious. That harm is a design-level risk rather than a measured outcome, because UDK and ATP denied all requests for the demographic data needed to test the models for bias, so no disparate-impact figure exists in the record. Oversight is thin: the Danish Data Protection Authority (Datatilsynet) can generally act only on complaints (GDPR Art. 57) with no proactive power, and because flagged people rarely learn an algorithm selected them, complaints are rare. UDK rejects the discrimination-by-design and social-scoring findings; no court has ruled.
Sources: amnestyinternationalalgorith2024, amnestyinternationaldanmark2024, bablai2024
Appears on: /domains/cases/denmark-udbetaling
EmpiricalIn judgment STS 1119/2025 of 11 September 2025, the Third Section of Spain's Supreme Court (Sala de lo Contencioso-Admin…
In judgment STS 1119/2025 of 11 September 2025, the Third Section of Spain's Supreme Court (Sala de lo Contencioso-Administrativo) ordered the government to give the transparency foundation Civio access to the source code of BOSCO, the software that determines eligibility for the electricity social bonus (bono social electrico). Applying the Transparency Law (Ley 19/2013) together with Article 42 of the EU Charter of Fundamental Rights and Article 105.b of the Spanish Constitution, the Court held that access to public information is a constitutional right and that neither intellectual property nor national security is an automatic shield, dismissing the government's secrecy claims as a 'mere risk' of eventual harm to be assessed case by case under a proportionality test. Civio and legal commentators describe an 'error multiplier': because BOSCO decides automatically and gives no reasons, one systematic error can propagate to thousands of eligible people at once. As of May 2026, roughly eight months after the ruling, the source code had still not been delivered and Civio had filed for judicial enforcement.
Sources: consejogeneraldelpoderjudici2025, fundacionciudadanacivio2025b, fundacionciudadanacivio2025a, derechoadministrativoyurbani2025, fundacionciudadanacivio2025c, fundacionciudadanacivio2026, freesoftwarefoundationeurope2026
Appears on: /domains/cases/spain-bosco
EmpiricalThe transparency foundation Civio documented, by reconstructing BOSCO's behaviour from partial technical specifications …
The transparency foundation Civio documented, by reconstructing BOSCO's behaviour from partial technical specifications and functional test cases, two systematic ways the software denied the electricity social bonus to people who qualified: when a pensioner ticked the 'pensioner' box the application could return an 'imposibilidad de calculo' (impossibility of calculation) error and be rejected without properly evaluating income; and large families, entitled to the bonus regardless of income, were denied whenever a household member withheld authorization to consult income data, although income was not a regulatory requirement for that category. After a 2017-2018 overhaul required all beneficiaries to re-apply by 31 December 2018, enrollment fell from roughly 2.4 to 2.5 million under the prior scheme to 1,111,958 as of January 2019 (later cited around 1.5 million), against an estimated 4.5 to 5.5 million eligible people, and more than half a million applicants were rejected. No audited per-decision error rate is public, because the source code and verification-test results were withheld; one academic analysis records errors in both directions, but the documented net effect is under-inclusion.
Sources: fundacionciudadanacivio2019, algorithmwatchnicolaskayserb2019, freesoftwarefoundationeurope2026, xatakaenriqueperez2024, rebootdemocracyjoseluismarti2025
Appears on: /domains/cases/spain-bosco
EmpiricalSerbia's Social Card (Socijalna karta) registry, given a statutory basis by the Law on the Social Card in force from 1 M…
Serbia's Social Card (Socijalna karta) registry, given a statutory basis by the Law on the Social Card in force from 1 March 2022 and financed in part by an 82.6 million euro World Bank public-sector loan, cross-links roughly 130 to 135 categories of data from other state registers to verify social-assistance eligibility and flag suspected undeclared income or assets. After the law, named sources report the caseload falling by a range of tens of thousands: government figures cited by Amnesty International show about 35,000 fewer recipients by August 2023, A11 counts at least 44,000 people having lost assistance by early 2024, and the UN Working Group on Business and Human Rights reported over 60,000 without assistance by October 2025. These are largely net caseload declines rather than audited counts of system-caused removals, and the government attributes part of the fall to a stronger economy. Roma are reported among the most affected because informal earnings are misclassified as income, but the registry records no ethnicity, so this is inferred rather than officially disaggregated. As of the latest reporting, Constitutional Court, World Bank Inspection Panel, and UN scrutiny were pending or active, with no court or panel yet ordering changes.
Sources: ainitiativeforeconomicandsoc2024, amnestyinternational2023b, contextthomsonreutersfoundat2023, unworkinggrouponbusinessandh2025, worldbankinspectionpanel2024, chinaceeinstitute2024
Appears on: /domains/cases/serbia-social-card
EmpiricalUnder Serbia's Social Card system, a removed beneficiary has 15 days to appeal and must wait three months to reapply reg…
Under Serbia's Social Card system, a removed beneficiary has 15 days to appeal and must wait three months to reapply regardless of changed circumstances, and removal letters frequently reference only unspecified data from the electronic database. A11's Request for Inspection to the World Bank Inspection Panel alleges that, because the system is semi-automated, social workers cannot correct errors recorded in it. Over roughly two years the Ministry processed more than 100,000 notifications of suspected income or asset increases, while beneficiaries filed only 361 appeals against Centers for Social Work rulings; because the two figures cover different populations, the gap illustrates how rarely flags were contested rather than a measured appeal rate. Documented misclassifications include a one-off funeral donation read as income and long-scrapped cars still counted as assets.
Sources: amnestyinternational2023b, ainitiativeforeconomicandsoc2024, worldbankinspectionpanel2024, chinaceeinstitute2024, contextthomsonreutersfoundat2023
Appears on: /domains/cases/serbia-social-card
EmpiricalSamagra Vedika, an entity-resolution system built by the Telangana government, decided welfare eligibility by matching r…
Samagra Vedika, an entity-resolution system built by the Telangana government, decided welfare eligibility by matching residents across thirty-plus government databases into a consolidated profile; between 2014 and 2019 more than 1.86 million ration cards were cancelled and 142,086 fresh applications were rejected without notice. Its core error was entity-resolution false-positive matching, in which a similarly-named third party's asset was attributed to the applicant and silently flipped the eligibility flag. After the Supreme Court of India ordered field re-verification in April 2022, a partial re-verification found roughly 7.5 percent wrongful rejection (at least 15,471 approved of 205,734 re-processed cases), a lower bound from an incomplete review; the system is proprietary and closed and an independent technical audit could not be completed, with no source code or accuracy data released. The government cited a self-reported 95 percent fraud-filtering efficiency, which measures spurious-application filtering rather than the wrongful-exclusion rate.
Sources: amnestyinternational2024c, tapasya2024, tusharvsharma2026, sumitjha2024, kumarsambhav2020, pulitzercenteraiaccountabili2024
Appears on: /domains/cases/india-samagra-vedika
EmpiricalUnder Samagra Vedika, exclusions were silent and there was no statutory route to contest an algorithmic decision, so the…
Under Samagra Vedika, exclusions were silent and there was no statutory route to contest an algorithmic decision, so the burden of proof fell on the excluded person: reporting describes officials who, though formally able to override the algorithm with evidence, deferred to it and declined to overturn its verdict, treating errors as backend technical issues. Documented individual harms include a 67-year-old widow denied rations for more than seven years after the system linked her deceased husband to a car owned by a similarly-named third person, and a family rejected for allegedly owning a four-wheeler that was declared eligible only after a Telangana High Court ruling. Corrections came through individual litigation and did not systematically feed back into the model, and the same entity-resolution technology was reused to issue new ration cards in 2024-2025.
Sources: tapasya2024, thereporterscollective2024, tusharvsharma2026, amnestyinternational2024c, sumitjha2024
Appears on: /domains/cases/india-samagra-vedika
EmpiricalIn A.M.C. v. Smith (No. 3:20-cv-00240, M.D. Tenn.), a federal court held after a five-day bench trial that Tennessee's D…
In A.M.C. v. Smith (No. 3:20-cv-00240, M.D. Tenn.), a federal court held after a five-day bench trial that Tennessee's Deloitte-built TEDS automated Medicaid eligibility system, operational statewide since March 19, 2019 for a program covering roughly 1.7 million residents, produced wrongful terminations, wrong-household assignments, and misleading or missing notices that violated the Medicaid Act, the Fourteenth Amendment's Due Process Clause, and the Americans with Disabilities Act; the 116-page opinion, issued August 26, 2024 by Judge Waverly D. Crenshaw Jr., ordered mediation before considering an injunction.
Sources: statescoopkeelyquinlan2024, stotlerhayesgroupllcerinsail2024, georgetownuniversitycenterfo2024, nationalhealthlawprogram2024
Appears on: /domains/cases/tennessee-tenncare-teds
EmpiricalThe UK government built its own AI meeting scribe for council caseworkers and piloted it through a cohort of 25 selected…
The UK government built its own AI meeting scribe for council caseworkers and piloted it through a cohort of 25 selected councils (22 active, more than 400 users) under one shared pooled-assurance record, then open-sourced it and adapted it to enlist around 500 housing and homelessness workers by June 2026; the cohort published a multi-council governance dataset but no transcription-accuracy or error-rate evaluation, and standard risk controls such as penetration testing and certification had not been completed on the alpha at pilot time.
Sources: localgovernmentassociation2025b, localgovernmentassociation2025a, ministryofhousing2026, trendall2026, incubatorforartificialintell2026
Appears on: /domains/cases/minute-local-ai
EmpiricalIndependent research by the Ada Lovelace Institute on AI transcription in social work, based on interviews with 39 socia…
Independent research by the Ada Lovelace Institute on AI transcription in social work, based on interviews with 39 social workers across 17 local authorities in England and Scotland, reported that local authorities focus their evaluations on efficiency rather than impact on people who draw on care and that perceptions of reliability and the need for human oversight vary significantly among workers; the research covers such tools sector-wide, not this tool specifically.
Sources: adalovelaceinstitute2026b, bruff2026, adalovelaceinstitute2026a
Appears on: /domains/cases/minute-local-ai
EmpiricalThe Ministry of Justice built an in-house AI transcription and summarisation copilot, Justice Transcribe, for probation …
The Ministry of Justice built an in-house AI transcription and summarisation copilot, Justice Transcribe, for probation staff in England and Wales, scaling it from a pilot to more than 1,000 officers in October 2025 and to every probation officer by June 2026, with official transparency data recording more than 800,000 supervision meetings summarised between 7 October 2025 and 2 June 2026; the reported time-savings are the ministry's own and rest on an operating assumption the department itself labels illustrative, and no transcription-accuracy rate, officer correction rate, or independent evaluation of the tool has been published.
Sources: justiceaiunit2026, ministryofjustice2025, ministryofjusticeandhmprison2025b, ministryofjusticeanddsit2025, ministryofjustice2026
Appears on: /domains/cases/justice-transcribe-probation
EmpiricalCopilot-written probation case records sit upstream of high-volume algorithmic risk assessment over the same record ecos…
Copilot-written probation case records sit upstream of high-volume algorithmic risk assessment over the same record ecosystem: reporting places the ministry's OASys-based reoffending-risk prediction at more than 1,300 people a day, drawing on probation and prison caseload systems and the Police National Computer, with a successor tool rolling out during 2026, and the ministry's own validation found lower predictive validity for all Black, Asian and Minority Ethnic groups for non-violent reoffending and for Black and Mixed ethnicity offenders for violent reoffending - a property of the downstream risk model, not the copilot; peer-reviewed commentary raises the erosion of professional judgment and the unresolved accountability for algorithm-influenced decisions as structural concerns, and no published source documents a named data pipeline from the copilot's output into the risk tools.
Sources: statewatch2025, phillips2026, nellis2026
Appears on: /domains/cases/justice-transcribe-probation
EmpiricalThe US Social Security Administration requires decision writers to run fully favorable disability decisions through its …
The US Social Security Administration requires decision writers to run fully favorable disability decisions through its in-house Insight verifier before issuance, with narrow documented exceptions, and the 2025 federal AI inventory records the tool computing 43 quality flags. In the agency's internal five-month study of roughly 50,000 appeals-level cases, reported through the 2019 Inspector General audit, analysts who used Insight logged about 0.9 errors per case against 0.7 for non-users, saw processing time fall about 4.7 days per case, and had about 12.6 percent of their cases returned for quality issues against 21.5 percent for non-users. These are internal, non-randomized comparisons among self-selected voluntary users, and the same audit found the agency stopped tracking performance after the first five months and could not determine any effect on remands.
Sources: ussocialsecurityadministrati2019, ussocialsecurityadministrati2026, engstrom2020
Appears on: /domains/cases/ssa-insight
EmpiricalIn a review issued April 30, 2026 (report 25-00153-47), the Department of Veterans Affairs Office of Inspector General f…
In a review issued April 30, 2026 (report 25-00153-47), the Department of Veterans Affairs Office of Inspector General found that at least 8,000 of an estimated 8,100 automated Dependency and Indemnity Compensation (survivor-benefit) granting decisions issued from September 2023 through August 2024 - nearly all - contained at least one legal or procedural deficiency, such as incomplete evidence summaries and omitted favorable findings, with most rating decisions listing only the death certificate as evidence. The OIG separately found that at least 2 percent of the decisions (at least 190) carried monetary-impact legal errors totaling at least 2.7 million dollars (2,727,764 dollars in questioned costs); the roughly 98 percent figure is the share with any legal or procedural defect, not the monetary-error rate. The system, phased in beginning May 2020, extracts data from scanned documents and applies predefined encoded rules to grant service-connected death claims end to end with no human involvement when the rules are met; the OIG describes it as rules-based automation and document extraction, not machine learning, and its figures are outcome statistics from a statistical sample rather than a per-interaction rate.
Sources: departmentofveteransaffairso2026, nieberg2026, weston2026
Appears on: /domains/cases/va-claims-automation
EmpiricalThe Office of Inspector General reported that VA's internal correction channels did not catch the automated survivor-ben…
The Office of Inspector General reported that VA's internal correction channels did not catch the automated survivor-benefit deficiencies and that the external audit was, empirically, the only channel that changed behavior. In April 2020 a VBA analyst reported through the internal defect-tracking system that automated decisions listed only the death certificate as evidence, and the Pension and Fiduciary Service closed the defect without action; the same deficiency was central to the 2026 findings, and VA removed the long-form guidance from its manual only in March 2025, immediately after the OIG's preliminary briefing - roughly five years later, and the OIG's full public report did not follow until 2026, roughly six years after the ticket. The OIG found the quality-review checklist for automated claims was less rigorous than the review traditional claims receive, and that the PACT Act section 701(b) modernization plan to Congress did not fully disclose that VBA grants these claims end to end without human intervention. Errors persisted as the program expanded: the VA Secretary announced expanded DIC automation in May 2025, and 20 additional automated decisions from September and October 2025 showed similar errors as of November 2025, with one recommendation still open and VBA concurring only in part.
Sources: departmentofveteransaffairso2026
Appears on: /domains/cases/va-claims-automation
EmpiricalIn Trelleborg, Sweden, the first municipality to fully automate social-assistance decisions, peer-reviewed analysis repo…
In Trelleborg, Sweden, the first municipality to fully automate social-assistance decisions, peer-reviewed analysis reports that about 30 percent of digital reapplications are decided entirely by rules-based software with no human review and about 85 percent receive at least partial automated handling; decision time on reapplications fell from roughly two days to under a minute, and a human caseworker re-enters the path only by exception, when a routing rule detects significantly changed circumstances, a missing activity plan or job-seeking documentation, or a complex or negative case. No error, override, exception-routing, or appeal-rate figures for the automated path have been published, so the fraction of automated decisions that ever reaches a human cannot be established from the record.
Sources: algorithmwatch2020b, ranerupandhenriksen2022, europeancommissionjointresea2021
Appears on: /domains/cases/trelleborg-rpa
EmpiricalThe City of Amsterdam spent roughly five years and an estimated EUR 535,000 building a deliberately fair, explainable we…
The City of Amsterdam spent roughly five years and an estimated EUR 535,000 building a deliberately fair, explainable welfare-fraud screening model with nearly every recommended pre-deployment safeguard in place - a bias audit, training-data reweighting that approximately equalized wrongful-flag rates on retrospective data, a data-protection assessment and a human-rights assessment, external and academic review, a citizen panel, and dual algorithm-register transparency - and discontinued it after a 2023 live pilot on nearly 1,600 applications. In the investigating journalists' analysis of aggregate data the city provided, the group disparities re-emerged inverted on the live pilot, now more likely to wrongly flag Dutch nationals, women, and applicants with children, with the tool flagging more applications than the analog process and no better than caseworkers at finding genuine cases. The Dutch national algorithm register records the deployment ending September 2023 and lists it out of use, and the responsible alderman announced the halt in November 2023.
Sources: braun2025, lighthousereportsa, algoritmeregisterdutchnation2023
Appears on: /domains/cases/amsterdam-slimme-check
EmpiricalIn a developer-reported randomised controlled trial of more than 1,000 adviser support requests, an adviser-facing benef…
In a developer-reported randomised controlled trial of more than 1,000 adviser support requests, an adviser-facing benefits copilot at Citizens Advice returned supervisor-checked answers in about four minutes, roughly half the previous response time, with about 80 percent of its drafts approved by supervisors without revision; these figures are reported by the tool's builders and have not been independently replicated.
Sources: varotsis2025, departmentforscience2025a, stanfordlegaldesignlabjustic
Appears on: /domains/cases/caddy-citizens-advice, /what-ai-can-do
EmpiricalAdvisers given access to the copilot were reported to be more than twice as likely to say they felt confident giving adv…
Advisers given access to the copilot were reported to be more than twice as likely to say they felt confident giving advice than a control group, a self-reported measure from post-call in-chat surveys rather than a client-outcome or accuracy measure.
Sources: varotsis2025, stanfordlegaldesignlab2025
Appears on: /domains/cases/caddy-citizens-advice, /what-ai-can-do
EmpiricalIn 2023 Singapore's GovTech began a whole-of-government retire-and-replace of its scripted Ask Jamie chatbots, embedded …
In 2023 Singapore's GovTech began a whole-of-government retire-and-replace of its scripted Ask Jamie chatbots, embedded since 2014 on 70-plus (a vendor case study claims 80) agency websites as independent per-agency answer engines, migrating government chatbots onto centrally provided large-language-model engines; the stated aim was to convert all 88 chatbots and retire the scripted engine by end 2023, the verified snapshot is 21 of 88 converted as of September 2023 (migration completion not independently documented), and by the VICA product page updated 29 April 2026 the successor platform hosts over 100 chatbots for 60-plus agencies at an average of over 800,000 monthly queries, figures that are all government self-reported.
Sources: hirdaramani2023, govtechsingapore2026, govtechsingapore2019
Appears on: /domains/cases/singapore-chatbot-fleet-refresh
EmpiricalThe Singapore government benefits-navigation surface is documented as scope-limited to information and estimates rather …
The Singapore government benefits-navigation surface is documented as scope-limited to information and estimates rather than adjudication: the Ministry of Finance Support For You Calculator turns self-declared inputs into benefit estimates that are explicitly estimates and not entitlement decisions, and the Chat.Gov.SG (Beta) explainer hosted on the SupportGoWhere domain states the assistant summarises information from official government websites and does not assess eligibility, make decisions, submit applications, or complete transactions, and warns users not to share personal or sensitive information.
Sources: publicservicedivisionsingapo2026, mustsharenews2024
Appears on: /domains/cases/singapore-chatbot-fleet-refresh
EmpiricalA June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported th…
A June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported that the IRS expanded its Automated Collection System chatbot and live-chat program and made live chat permanent while having no performance measures for it, despite a Taxpayer First Act requirement for metrics and benchmarks, and that management's claim the bots reduced telephone demand could not be substantiated; the statistical reports the IRS did collect were deemed unreliable, in one instance showing a single assistor apparently working 603 chats at once against a systemic cap of three, attributed partly to a miscalculated handle-time metric the vendor had not resolved as of December 2025.
Sources: treasuryinspectorgeneralfort2026, bracken2026, bramwell2026
Appears on: /domains/cases/irs-acs-chatbots
EmpiricalIn the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple chats concu…
In the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple chats concurrently and 12 of those 24 had at least one authenticated chat open while working another, which TIGTA reported as raising the risk of disclosing taxpayer information to the wrong taxpayer; the audit also reported 635,684 resolution codes against 613,056 chats (a mismatch management knew of but did not investigate) and, in March 2025 hand-testing, 14% of chatbot process flows deficient and 83% of tested keywords unrecognized or insufficient, with the figures drawn from a nonprobability sample and data the audit itself characterized as unreliable and not projectable to the full assistor population.
Sources: treasuryinspectorgeneralfort2026, bramwell2026, cohn2026
Appears on: /domains/cases/irs-acs-chatbots
EmpiricalIn a spring-2025 randomized pilot inside a large consumer EBT app, the vendor reports that 53% of eligible SNAP recipien…
In a spring-2025 randomized pilot inside a large consumer EBT app, the vendor reports that 53% of eligible SNAP recipients took up in-app AI help for missed deposits and that treated users were restored faster and more often in the same month than a control group, with every AI dead-end escalated to a named human; all outcome figures are vendor-published and the effect magnitudes were not disclosed.
Sources: propelincpropelinsights2025b, guarino2025a
Appears on: /domains/cases/propel-snap-assistant
EmpiricalBy the vendor's own account of the design, the assistant grounds on a state-verified deposit record it reads but does no…
By the vendor's own account of the design, the assistant grounds on a state-verified deposit record it reads but does not write to, and steers recipients to act on the state system of record rather than acting for them.
Sources: propelincpropelinsights2025b
Appears on: /domains/cases/propel-snap-assistant
EmpiricalBetween 2019 and 2025 more than 70% (about 73% per its ten-year retrospective) of California's online SNAP applications …
Between 2019 and 2025 more than 70% (about 73% per its ten-year retrospective) of California's online SNAP applications were submitted through GetCalFresh, a deterministic, structured-workflow application assister built and operated by the nonprofit Code for America, which reports helping 6.2 million people obtain more than $12.8 billion in food benefits from 2017 to 2025 (organization-published figures that are not independently audited); the node made no eligibility determinations, and in 2024 and 2025 the California Department of Social Services coordinated a dated, phased transfer of its functions into the state-owned BenefitsCal portal.
Sources: codeforamerica2024b, codeforamerica2025a, californiadepartmentofsocial2025
Appears on: /domains/cases/getcalfresh, /what-ai-can-do
EmpiricalA randomized controlled trial of roughly 65,000 Los Angeles GetCalFresh applicants (Giannella, Homonoff, Rino, and Somer…
A randomized controlled trial of roughly 65,000 Los Angeles GetCalFresh applicants (Giannella, Homonoff, Rino, and Somerville, American Economic Journal: Economic Policy 16(4), 2024) found that access to applicant-initiated flexible interviews increased SNAP approvals by about 6 percentage points, doubled early approvals, and raised long-term participation by over 2 percentage points, identifying the intake interview as a key procedural-denial barrier; Code for America separately reported an in-house experiment lifting renewal-form submissions among prior non-responders from about 1.5% to roughly 12% (organization-published, without sample sizes or confidence intervals).
Sources: giannella2024, codeforamerica2021, codeforamerica2024a
Appears on: /domains/cases/getcalfresh
EmpiricalIn June 2024 the board of Benefits Data Trust, a Philadelphia benefits-navigation nonprofit that reported helping more t…
In June 2024 the board of Benefits Data Trust, a Philadelphia benefits-navigation nonprofit that reported helping more than 120,000 people access about $182 million in benefits in 2023, voted unanimously to wind the organization down within a self-imposed 60-day window, citing only 'a perfect storm of circumstances'; the organization closed on August 24, 2024, laying off 273 employees, despite roughly $12 million in unrestricted reserves at the end of 2023 and about $32 million in projected 2024 revenue.
Sources: brubaker2024a, brubaker2024b, wink2024, mosbruckergarza2024
Appears on: /domains/cases/benefits-data-trust-winddown
EmpiricalThe closure left active government partnerships without a designated successor, including a Pennsylvania Department of A…
The closure left active government partnerships without a designated successor, including a Pennsylvania Department of Aging workload of nearly 48,000 applications from 27,018 households in the final year and a Philadelphia BenePhilly call-center contract the organization was reported to be exceeding through mid-2024; the navigation function fragmented to higher-friction channels, with the work redistributed across partner agencies and a subcontractor and referral waits reported as several months, which a Pew analyst described as a 'cascading effect.'
Sources: brubaker2024d, burnley2024, mosbruckergarza2024
Appears on: /domains/cases/benefits-data-trust-winddown
EmpiricalLondon's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately governed syste…
London's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately governed systems - CHAIN street-outreach contacts, In-Form charity casework, and H-CLIC borough statutory applications - into a single rough-sleeping journey per person that is read, in aggregate form only, across all 33 London local authorities; the tool makes no individual-level determinations, and after the build vendor's data-processor contract ended on 2 February 2024 the Greater London Authority contracted Homeless Link, which also operates the CHAIN source system, for its ongoing hosting, management, and maintenance.
Sources: techuk2024, loti2023, londonofficeoftechnologyandi2023
Appears on: /domains/cases/london-rough-sleeping-sit
EmpiricalThe Strategic Insights Tool's matcher accepts an association only above an 85% probability threshold chosen to minimise …
The Strategic Insights Tool's matcher accepts an association only above an 85% probability threshold chosen to minimise false positives, and the project's own Phase 2 Data Protection Impact Assessment reports 91% recall - conceding that roughly 9 in 100 true cross-system matches are missed so that 'numbers subsequently appear lower in places where they should be higher' and that recall varies as new data of varying quality is ingested; no false-positive rate is published, the accuracy figures are self-reported by the delivery team, and no independent evaluation of the tool's decision impact exists.
Sources: loti2023, lotiannahumplebyandfacultyja2025
Appears on: /domains/cases/london-rough-sleeping-sit
EmpiricalSan Jose's vehicle-mounted computer-vision pilot, described by city officials and national housing advocates as the firs…
San Jose's vehicle-mounted computer-vision pilot, described by city officials and national housing advocates as the first US experiment training AI to recognize tents and lived-in vehicles, reported sharply class-asymmetric accuracy in the city's own staff-ground-truthed evaluation — 97% for potholes and 88% for trash, but only 70% for RVs (unable to distinguish a lived-in RV from an empty one) and 12.5% for lived-in vehicles, with a March 2024 official interview bracketing the habitation figures at 70–75% for RVs and 10–15% for lived-in cars against a 70% goal; no detection ever generated an operational dispatch, and after investigative exposure and structured engagement the city removed every habitation-detection use case, its March 2025 status report declining to recommend implementing AI object detection in city operations at this time.
Sources: feathers2024, cityofsanjoseinformationtech2025, usdepartmentoftransportation2025
Appears on: /domains/cases/san-jose-encampment-detection
EmpiricalThe pilot's published data-usage protocol declares that the footage cannot be actively monitored for law-enforcement pur…
The pilot's published data-usage protocol declares that the footage cannot be actively monitored for law-enforcement purposes while preserving a police request path to it — verbatim, 'Law enforcement may request access to previously stored footage. Law enforcement is not actively monitoring any data collected' — and requires de-identification or deletion within one month; the CIO stated data was not shared with police during the pilot, yet public-records reporting documented that one vendor's system ran optical character recognition of license plate numbers despite the city's no-identification claim, so the declared authority rule and the feasible data flows diverged, a gap surfaced by journalists rather than by any standing audit, and the no-law-enforcement-use clause is city protocol language rather than statute.
Sources: feathers2024, cityofsanjoseinformationtech2024, varian2024
Appears on: /domains/cases/san-jose-encampment-detection
EmpiricalThe Los Angeles Coordinated Entry System replaced the VI-SPDAT survey for single adults with the Los Angeles Housing Ass…
The Los Angeles Coordinated Entry System replaced the VI-SPDAT survey for single adults with the Los Angeles Housing Assessment Tool, a 19-item self-report score whose weights were derived by a regression on 71,747 historical assessments linked to county records; where the CESTTRR research estimated the VI-SPDAT scored near chance (AUC 0.54) with racial false-negative gaps up to 8.5 percentage points, the equity-adjusted successor was deliberately traded down in overall accuracy (AUC 0.60, from an accuracy-only 0.64) to close those gaps to under one percentage point, and every such figure is a pre-deployment estimate on 2015 to 2018 held-out data rather than an observed post-launch outcome.
Sources: rice2023, losangeleshomelessservicesau2025a
Appears on: /domains/cases/lahsa-triage-revision
EmpiricalDuring the dual-tool transition the two instruments' PSH-consideration thresholds were 8-plus on the VI-SPDAT and 17-plu…
During the dual-tool transition the two instruments' PSH-consideration thresholds were 8-plus on the VI-SPDAT and 17-plus on the LA HAT, and by LAHSA's account initial quantitative data and provider feedback showed participants were more likely to obtain an eligible score under the VI-SPDAT, so direct-service providers opted to administer it, a trend LAHSA states 'perpetuated the racial bias of the VI-SPDAT in the System'; on April 22, 2026 the CES Policy Council lowered the LA HAT threshold to 12-plus, ruled the most recent LA HAT score supersedes a coexisting VI-SPDAT score, and forced deactivation of new VI-SPDAT completions (for LA HAT-access programs on May 1, 2026 and system-wide on June 30, 2026), though LAHSA has not released the underlying eligibility-rate numbers.
Sources: losangeleshomelessservicesau2025a, losangeleshomelessservicesau2026b, losangeleshomelessservicesau2026a
Appears on: /domains/cases/lahsa-triage-revision
EmpiricalIn a registered randomized controlled trial of 1,263 imminent-risk applicants (514 treatment, 749 control) run by the Un…
In a registered randomized controlled trial of 1,263 imminent-risk applicants (514 treatment, 749 control) run by the University of Notre Dame's evaluation lab, households offered flexible emergency financial assistance averaging about 2,000 dollars, typically one to two months of back rent, through Santa Clara County's homelessness-prevention system were reported 81 percent less likely to become homeless within six months and 73 percent within twelve months; the peer-reviewed article's abstract states the assistance reduced homelessness by 3.8 percentage points from a 4.1 percent base rate, and the researchers conservatively estimated 2.47 dollars in community benefits per net dollar spent.
Sources: phillipsandsullivan2025, universityofnotredamenews2023, phillipsandsullivan2021
Appears on: /domains/cases/santa-clara-prevention
EmpiricalBecause becoming homeless is statistically rare even among at-risk applicants - about 96 percent of the trial's control …
Because becoming homeless is statistically rare even among at-risk applicants - about 96 percent of the trial's control group never became homeless without assistance - the program's own co-author cautions that prevention resources can flow to households that would have stayed housed anyway, making screening precision on a low base rate the binding constraint; as of February 2026 the model is being replicated across about ten heterogeneous US jurisdictions under a 77-million-dollar initiative, with the same evaluation lab as the common evidence partner assessing each site.
Sources: kendall2026a, phillipsandsullivan2025, destinationhome2026b, universityofnotredamenews2026
Appears on: /domains/cases/santa-clara-prevention
EmpiricalAt the Calgary Drop-In Centre, a University of Calgary engineering group and the NGO shelter operator built deliberately…
At the Calgary Drop-In Centre, a University of Calgary engineering group and the NGO shelter operator built deliberately interpretable screening for chronic and episodic shelter use - explicit stay-count thresholds (for example 81 or more stays in a 90-day window) and database-queryable rules derived from the shelter's own administrative records, reported to flag candidate clients at a median of about 98 days versus 285 days under the Government of Canada definition and 365 under the Alberta definition - and, rather than surface a risk score, deployed a co-designed data-navigation interface that shows frontline staff raw client histories; no fetched source confirms the thresholds running as an automated production screener, and the deployed, studied artifact is the raw-history interface.
Sources: messier2021, arulesearchframeworkfortheea2022, masrani2025
Appears on: /domains/cases/calgary-drop-in-shelter-ml
EmpiricalAcross a 2022 to 2024 embedded deployment study of the interface (16 staff across 7 role categories; 29.5 hours of quali…
Across a 2022 to 2024 embedded deployment study of the interface (16 staff across 7 role categories; 29.5 hours of qualitative data; five committee observations; three deployed versions), the participant-research team documented a stakes-dependent 'data-outsourcing continuum': staff were reluctant to outsource high-stakes barring decisions, treating the data as a starting point for collaborative discussion, while reporting more willingness to accept automated data-driven recommendations for lower-stakes housing triage; the finding is the staff's own articulated practice rather than a measured override or agreement rate, all deployment evidence is authored by the embedded research team, and no independent evaluation, usage logs, or decision volumes are published.
Sources: masrani2025, thehumanbehindthedatareflect2023
Appears on: /domains/cases/calgary-drop-in-shelter-ml
EmpiricalAn AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather t…
An AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather than callers, expanding measured review from the under-3% of calls that had been reviewed by hand toward nearly all of them; a peer-reviewed reliability study of 476 labeled calls reported agreement with human ratings at 98 percent of human interrater agreement for detecting any risk assessment, with average F1 of about 0.86 at call level and 0.66 at statement level, and its authors include four holders of equity in the vendor.
Sources: imel2024, aguilar2023, nihreporternationalinstitute2025
Appears on: /domains/cases/lyssn-protocall-988
EmpiricalThe registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on October 31, 202…
The registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on October 31, 2025, but as of mid-2026 no results were posted to the trial registry or found in the peer-reviewed literature and participant-level data were marked unavailable for proprietary reasons, so reported counselor-skill-improvement effects remain vendor claims pending independent publication.
Sources: clinicaltrialsgovusnationall2026, lyssn2026
Appears on: /domains/cases/lyssn-protocall-988
EmpiricalGaggle's student-communication safety monitoring, used by roughly 1,500 US districts covering about 6 million students a…
Gaggle's student-communication safety monitoring, used by roughly 1,500 US districts covering about 6 million students as of a March 2025 AP and Seattle Times investigation, scans school-issued accounts around the clock and routes flags through a multi-hop chain (a machine flag, an off-site vendor reviewer, district safety staff, and, for imminent-danger after-hours alerts, occasional police welfare checks); in Vancouver Public Schools nearly 2,200 students (about 10% of enrollment) triggered alerts in one year, in Lawrence USD 497 more than 1,200 incidents were logged in ten months with about two-thirds deemed nonissues by officials (a figure the plaintiffs drew from district records), and the archive of flagged documents was accidentally released to reporters as nearly 3,500 unredacted files through unprotected links, while a 2023 RAND review found only scant evidence of either benefit or risk and the vendor publishes no accuracy figures.
Sources: bryanandlurye2025, associatedpress2025, lawrencejournalworld2025
Appears on: /domains/cases/gaggle-school-monitoring
EmpiricalAfter nine Lawrence, Kansas students sued their district in early August 2025 over its use of AI communication monitorin…
After nine Lawrence, Kansas students sued their district in early August 2025 over its use of AI communication monitoring, court filings revealed the district had ceased using Gaggle mid-litigation and substituted a different monitoring vendor with no board vote or public disclosure — surfacing only as a line in a check register — and the plaintiffs' amended complaint argued the swap does not moot the case because the core practice of suspicionless scanning, flagging, and seizure of student speech continues; on April 10, 2026 a federal judge found the district violated the Kansas Open Records Act in withholding the substitution and phase-out records, and on June 4, 2026 ordered it to pay the students' attorney fees, characterizing the conduct as drawn out, hollow and perplexing, with a jury trial on the surviving constitutional claims set for January 2027.
Sources: heimsoth2025, heimsoth2026a, heimsoth2026b
Appears on: /domains/cases/gaggle-school-monitoring
EmpiricalIn March 2026 the UK Parliamentary and Health Service Ombudsman partly upheld a complaint that an NHS mental health trus…
In March 2026 the UK Parliamentary and Health Service Ombudsman partly upheld a complaint that an NHS mental health trust installed camera-based, contact-free bedroom monitoring on a psychiatric ward without seeking a patient's consent, gave her no information about it, and did not switch it off when she asked; the case documentation and investigative reporting describe an internal clinical evaluation that the vendor is reported to have authored the business case for and shaped, a rebrand of the vendor during a statutory inquiry, and an open data-protection investigation, while the tool's own outcome-reduction figures are vendor claims contested by a campaign-linked meta-analysis and its adoption share across NHS mental health trusts is reported only as a contested range.
Sources: parliamentaryandhealthservic2026, williamson2026a, williamson2026b, nationalsurvivorusernetwork2025, stopoxevision2026
Appears on: /domains/cases/oxevision-nhs-wards
EmpiricalThe ombudsman's report on the case (decision 27 March 2026) found the trust did not seek or revisit the patient's consen…
The ombudsman's report on the case (decision 27 March 2026) found the trust did not seek or revisit the patient's consent for the bedroom monitoring, did not turn the camera off when she asked, gave her no information about it, and kept no record of how staff used it, and that even the trust's revised 2025 procedure still permits overriding a capacitous patient's refusal on clinically-safe grounds with multidisciplinary-team approval; on the separate question of over-reliance the ombudsman found on balance, cross-referencing observation charts, a nurse-adviser review and door key-card data, that in-person observations had continued and did not uphold that part of the complaint.
Sources: parliamentaryandhealthservic2026
Appears on: /domains/cases/oxevision-nhs-wards
EmpiricalIn 2025 ODMAP's pre-set county thresholds - a rolling 24-hour count against a threshold each agency sets or accepts, rec…
In 2025 ODMAP's pre-set county thresholds - a rolling 24-hour count against a threshold each agency sets or accepts, recommended by the system as two standard deviations above the county's own previous 90-day mean, a deterministic rule rather than a machine-learning model - fired 74,805 advisory spike-alert notifications from 498,003 suspected, unconfirmed overdose events that only about 1,362 of its 5,605 approved agencies actually submitted; these figures are self-published by the program in its own annual report and manuals, and ODMAP states its data are suspected, incomplete, not a system of record, and should not be generalized beyond participating agencies.
Sources: washingtonbaltimorehidta2025b, washingtonbaltimorehidta2026c, washingtonbaltimorehidta2025a
Appears on: /domains/cases/odmap-overdose-spike-alerts
EmpiricalODMAP's shared overdose store is housed inside a federal drug-enforcement program, and its operating policies both state…
ODMAP's shared overdose store is housed inside a federal drug-enforcement program, and its operating policies both state that ODMAP is neither an intelligence sharing database nor a pointer index records system and grant the host permission to use the data as the HIDTA sees fit, including combining it with other databases it manages for law enforcement and public health products; a 2024 peer-reviewed stakeholder study documented divergent public-health versus public-safety data-privacy standards, and a 2025 peer-reviewed analysis argues the integration risks racialized surveillance and criminalization of people who experience overdose, a contested scholarly critique of the link structure rather than a documented misuse incident.
Sources: washingtonbaltimorehidta2022, syvertsen2025, allen2024
Appears on: /domains/cases/odmap-overdose-spike-alerts
EmpiricalThe Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospe…
The Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospectively across five hospitals of an academic health system covering 590,736 monitored patients — the largest prospective study of an ML sepsis system on record. Its central finding was conditional on the human loop: sepsis patients whose alert was evaluated and confirmed by a provider within three hours had a 3.3 percentage-point absolute and 18.7 percent relative adjusted reduction in in-hospital mortality, with less organ failure and shorter stays, while the alert on its own did not; a companion study found provider uptake varied with experience, unit culture, and alert context.
Sources: adams2022a, henry2022a
Appears on: /domains/cases/johns-hopkins-trews, /pan-lab
EmpiricalThe TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: it was bui…
The TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: it was built at the deploying institution and commercialized through a company founded by its principal investigator, and confirmation-associated benefit is an observational association rather than a randomized effect of the algorithm — providers who engaged with alerts may differ from those who did not in ways the adjustment does not capture. The strongest numbers in the record therefore come from the party with the strongest interest in them, and no independent replication of the mortality effect had been published.
Sources: adams2022a
Appears on: /domains/cases/johns-hopkins-trews, /pan-lab
EmpiricalThe Advance Alert Monitor is an in-hospital deterioration model running around the clock across 21 hospitals of an integ…
The Advance Alert Monitor is an in-hospital deterioration model running around the clock across 21 hospitals of an integrated health system, scoring inpatients hourly and firing roughly twelve hours before predicted deterioration; a 2020 New England Journal of Medicine evaluation associated its alert-driven rapid-response workflow with lower mortality. Its defining feature is where the alert goes: not to the bedside, but to a dedicated regional tier of critical-care virtual quality nurse consultants who screen every alert around the clock, work up the chart, and only then escalate to the on-site rapid-response team — so the measured benefit is priced against the whole two-tier staffing topology, not the model alone.
Sources: escobar2020a, thekaiserpermanentenorthernc2022
Appears on: /domains/cases/kaiser-aam-deterioration, /pan-lab
EmpiricalSepsis Watch is a deep-learning sepsis-detection system scoring every emergency-department patient every five minutes ov…
Sepsis Watch is a deep-learning sepsis-detection system scoring every emergency-department patient every five minutes over 86 variables, deployed at an academic hospital under a registered clinical trial, with alerts fronted by rapid-response-team nurses who track treatment-bundle completion on three- and six-hour timers. Its structural fault line is an authority split: the operator who receives the alert (the nurse) is not the operator empowered to act on it (the physician who holds treatment authority), so the correction runs through a peer-persuasion edge. An independent ethnography found the system worked because nurses performed hidden repair work — mediating the professional hierarchy and doing the emotional labor of communicating a risk score upward — labor that was structurally necessary, largely invisible to the deployment's formal description, and undervalued.
Sources: sendak2020a, elish2020
Appears on: /domains/cases/duke-sepsis-watch, /pan-lab
EmpiricalA widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and s…
A widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and switched on across hundreds of hospitals was externally validated in 2021 across 38,455 hospitalizations at an academic health system: it achieved an area under the curve of 0.63, identified only 33 percent of sepsis cases, and had a positive predictive value of about 12 percent, generating roughly 109 alerts for every true sepsis case — a real-world performance the vendor had not fully examined before selling the model, and which an investigation attributed in part to undisclosed features such as antibiotic-order data that inflated internal validation.
Sources: wong2021c, statnews2021
Appears on: /domains/cases/epic-sepsis-michigan, /pan-lab
EmpiricalAfter external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, …
After external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, and reducing its reliance on antibiotic-order features. A 2026 multicenter prospective validation of the updated model across 227,091 encounters reported an area under the curve of 0.82 to 0.92 with positive predictive value of 0.13 to 0.26 and substantial between-site variability, and its authors urged local validation and alert-silencing strategies rather than trusting the model out of the box — a correction that arrived only after independent scrutiny of a model that had already been deployed at scale behind a corporate firewall shielding it from outside inspection.
Sources: statnews2022, wong2026a
Appears on: /domains/cases/epic-sepsis-michigan, /pan-lab
EmpiricalThe largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7…
The largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7,260 physicians and 2,576,627 patient encounters over fourteen months, with roughly 16,000 hours of documentation time saved and sustained physician support measured along the way. The system records the visit and drafts the clinical note; the clinician edits and signs, and the model-to-record write is gated both by that clinician review and by a standing internal quality-assurance program over the AI output — a real subsystem with a real cost, because the drafted note becomes a permanent record that later clinicians and later tools read as fact.
Sources: tierney2024a, tierney2025a
Appears on: /domains/cases/kaiser-tpmg-scribe, /pan-lab
EmpiricalThe scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and should be …
The scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and should be read as that system's dashboard rather than a guarantee of the product class: a multisite study of 8,581 clinicians across five health systems found more modest effects — on the order of 13 to 16 fewer minutes per day with no meaningful after-hours relief — and a validated per-note evaluation found hallucinations in about 31 percent of ambient-generated notes under structured review, versus about 20 percent of physician-written gold-standard notes, making ambient notes more thorough but less accurate. The clinician review and quality-assurance program are the controls that stand between that error rate and a contaminated permanent record.
Sources: rotenstein2026a, palm2025a
Appears on: /domains/cases/kaiser-tpmg-scribe, /pan-lab
EmpiricalThe strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized t…
The strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized trial of an ambient scribe across 66 practitioners and 71,487 notes (38 percent AI-generated), which found work exhaustion significantly reduced, professional fulfillment unchanged (a recorded null), roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the larger first-party deployment reports, this is a randomized estimate of the well-being and time effects — though it measured practitioner well-being and time, not per-note error rates.
Sources: afshar2025a
Appears on: /domains/cases/uw-health-abridge-scribe, /pan-lab
EmpiricalThe same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monito…
The same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monitoring of ambient AI in production — the rare case where the organization-side monitoring function exists as a citable, designed subsystem rather than an assumed practice. That monitoring is the org's stated answer to a documented system-level risk of the technology: a coding arms race, in which better AI documentation raises coding intensity, payers recalibrate in response, and clinician attestation liability grows — so the improved coding accuracy the trial measured sits next door to an upcoding pressure the monitoring is meant to watch.
Sources: afshar2025b, dai2025a
Appears on: /domains/cases/uw-health-abridge-scribe, /pan-lab
EmpiricalA peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note time per ap…
A peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note time per appointment reduced (6.2 to 5.3 minutes) and NASA-TLX cognitive load reduced — but the burnout change (42.1 to 35.1 percent) was not statistically significant, the domain's honest null bound of cognitive-load relief without a demonstrated burnout effect.
Sources: stults2025a
Appears on: /domains/cases/sutter-ambient-scribe, /pan-lab
EmpiricalIn the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians reported impr…
In the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians reported improved satisfaction against 36.4 percent of medical specialists — the same tool, in the same system, under the same workflow, helping one operator class and largely failing another, so any uniform service term overstates the effect for the group it helps least.
Sources: stults2025a
Appears on: /domains/cases/sutter-ambient-scribe, /pan-lab
EmpiricalCompany-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 de…
Company-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 developers at three enterprises found a pooled 26.08 percent increase in completed tasks, with gains concentrated among less-experienced developers. An independent randomized study of 16 experienced open-source maintainers on 246 tasks in familiar repositories bounded the expert tail from the other direction: those developers were about 19 percent slower with the AI while believing themselves about 20 percent faster — a measured perception-reality gap that means a uniform productivity number overstates the effect for senior engineers.
Sources: cui2025a, becker2025a
Appears on: /domains/cases/msft-accenture-copilot, /pan-lab
EmpiricalIndividual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry…
Individual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry research program measured a roughly 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability for every 25 percent increase in AI adoption, evidence that the churn the assistant adds must be absorbed by code-review and testing gates or the individual speed-up degrades the organization's delivery performance.
Sources: googleclouddora2024
Appears on: /domains/cases/msft-accenture-copilot, /pan-lab
EmpiricalAn in-house machine-learning code-completion system built, deployed, and measured by a company's own platform organizati…
An in-house machine-learning code-completion system built, deployed, and measured by a company's own platform organization for more than 10,000 internal developers reported, against a control group, a 25 to 34 percent suggestion-acceptance rate, a 6 percent reduction in coding iteration time versus control, and 3 percent of new code characters coming from the model at the time of measurement. The measuring party, the building party, and the deploying party were the same organization, and the numbers were published as an engineering-blog self-report rather than a peer-reviewed or independent evaluation.
Sources: tabachnyk2022
Appears on: /domains/cases/google-internal-completion, /pan-lab
EmpiricalThe same company's cross-industry research program reported that AI-assisted software development amplifies an organizat…
The same company's cross-industry research program reported that AI-assisted software development amplifies an organization's existing strengths and weaknesses rather than substituting for them, with policy clarity and platform investment identified as the levers that determine whether AI adoption improves or degrades delivery — evidence that the individual coding gains do not compose to organization-level outcomes on their own, and that the deploying organization's existing gates and platform quality are what decide the result.
Sources: googleclouddora2025
Appears on: /domains/cases/google-internal-completion, /pan-lab
EmpiricalA regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a co…
A regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, publishing its own measurement of the rollout. The bank's engineers reported productivity and code-quality improvements — and recorded the security impact as explicitly inconclusive, a real gating decision taken and documented under uncertainty rather than resolved by assertion, with the honestly recorded unknown carried forward into the scaled deployment.
Sources: chatterjee2024a, theregister2024
Appears on: /domains/cases/anz-bank-copilot, /pan-lab
EmpiricalWhat the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of cod…
What the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of code generated by a widely used assistant found that about 40 percent of generated programs contained vulnerabilities across scenarios spanning the CWE top-25 weaknesses, and separate research documents developers accepting insecure suggestions with overconfidence — so the security unknown a deployment carries forward unresolved sits against a class-level literature in which insecure generation is common.
Sources: pearce2022a
Appears on: /domains/cases/anz-bank-copilot, /pan-lab
EmpiricalA mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more tha…
A mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more than 400 developers, publishing acceptance telemetry (a 33 percent suggestion-acceptance rate, with 20 percent of suggested lines accepted), a 72 percent satisfaction figure, documented per-language variation, and stated limitations. Its evaluation instrument is acceptance-rate telemetry — which the productivity literature identifies as the measure most correlated with perceived productivity rather than outcome, and perception is measured to be miscalibrated for experienced developers, so acceptance telemetry captures adoption feel, not delivered output.
Sources: bakal2025a, ziegler2024a
Appears on: /domains/cases/zoominfo-copilot, /pan-lab
EmpiricalThe deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknown, one ste…
The deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknown, one step less honest than a deployment that runs a security check and records the result as inconclusive, because an absence no one has written down is not a governed object and cannot be carried forward or resolved. The value of the case is the documentation quality of an ordinary, competent adoption — phase gates, telemetry definitions, per-language deltas, and stated limitations by the deployer itself — with the missing security question priced as the one thing even that documentation did not name.
Sources: bakal2025a
Appears on: /domains/cases/zoominfo-copilot, /pan-lab
EmpiricalA global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering p…
A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Every one of those numbers is a vendor-and-customer self-report with no independent audit — which is itself the honest structure of the domain, because a peer-reviewed deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, and the alert-volume reduction the vendor advertises is precisely the lever a regulator scrutinizing an under-monitoring risk would question.
Sources: googlecloud2023
Appears on: /domains/cases/hsbc-aml, /pan-lab
EmpiricalTwo structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dom…
Two structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dominated by the false-alarm rate rather than by accuracy, so at realistic prevalence a threshold change moves the burden of alerts rather than the truth of them (the base-rate fallacy). And the labels the model learns from are the investigators' own dispositions: only a small set of flagged transactions is ever verified, and models are retrained on the analysts' calls, so a rise in 'confirmed' activity is partly a measure of what the system taught its reviewers to confirm rather than an independent ground truth (the label-feedback loop).
Sources: axelsson2000a, dalpozzolo2018a
Appears on: /domains/cases/hsbc-aml, /pan-lab
EmpiricalA neobank's fraud algorithms — triggered heavily by pandemic-era government benefit deposits — froze and closed the acco…
A neobank's fraud algorithms — triggered heavily by pandemic-era government benefit deposits — froze and closed the accounts of legitimate customers at scale, holding their balances for thirty to more than ninety days, and the company admitted some of the closures were mistakes. The false-positive tail here lands on real people as immediate hardship, concentrated among benefit-deposit recipients and low-balance households for whom a frozen account means no access to funds for weeks.
Sources: kessler2021
Appears on: /domains/cases/chime-fraud, /pan-lab
EmpiricalA 2024 federal consent order priced the downstream operational failure rather than the model: thousands of consumers wai…
A 2024 federal consent order priced the downstream operational failure rather than the model: thousands of consumers waited weeks to months for their balances after account closure, and the order imposed a 3.25 million dollar civil penalty plus at least 1.3 million dollars in consumer redress for the delayed refunds. The harm ran through three stages inside the organization's control — the scoring model's false positives, the operations backlog that turned a freeze into months without funds, and the refund process whose delay drew the regulator — and the enforcement attached to the last stage, the backlog, not to the model that started it.
Sources: consumerfinancialprotectionb2024
Appears on: /domains/cases/chime-fraud, /pan-lab
EmpiricalA Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive r…
A Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive rate — a measured pre-machine-learning baseline whose badness is the most credible datum in the record, since a 99.5 percent false-positive rate is not a marketing claim. The vendor-published rollout of a deep-learning engine scoring transactions in real time (under 300 milliseconds) claims false positives cut by about 60 percent and true-positive detection raised by about 50 percent; those figures are an organization-named, trade-press-covered vendor case study, entered here as claimed magnitudes against that legacy baseline because they were not independently audited.
Sources: teradata2017, groenfeldt2017
Appears on: /domains/cases/danske-fraud, /pan-lab
EmpiricalThe same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims o…
The same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims of authorized-push-payment scams in the regulator's bank-by-bank performance data — better detection and worse victim-outcome performance coexisting in one organization. And it was a rule, not a model, that moved the institutional behavior: the regulator's mandatory-reimbursement regime raised sector reimbursement from roughly two-thirds to about 89 percent, demonstrating that detection quality and the justice of the disposition are different levers held by different actors, and that the victim-outcome lever is a regulatory rule rather than a better classifier.
Sources: ukpaymentsystemsregulator2023
Appears on: /domains/cases/danske-fraud, /pan-lab
EmpiricalAn internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role…
An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on ten years of the company's own hiring decisions, a period whose hires were predominantly male. The models learned that history: they penalized the word 'women's' and downgraded graduates of women's colleges, reading gender proxies as negative signal. The team patched the identified terms but concluded that term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern rather than the words, and the company scrapped the project around 2017; per the company, recruiters saw the tool's recommendations but it was never used as a sole ranking.
Sources: dastin2018
Appears on: /domains/cases/amazon-resume-engine, /pan-lab
EmpiricalTraining a screener on an organization's past hiring decisions imports the past's selection function: research on hiring…
Training a screener on an organization's past hiring decisions imports the past's selection function: research on hiring as exploration finds that models trained on prior hires raise hire rates but replicate historical selection, and that a screener which values exploration rather than only exploitation breaks that lock-in loop. Two governance lessons follow — the patch lever has a documented ceiling, since removing named proxies does not remove a learned correlation, and abandonment can itself be a governance outcome, taken here before any external harm was documented rather than after an adjudication.
Sources: li2020a
Appears on: /domains/cases/amazon-resume-engine, /pan-lab
EmpiricalAn applicant-tracking platform whose AI screening and recommendation features operate inside thousands of employers' hir…
An applicant-tracking platform whose AI screening and recommendation features operate inside thousands of employers' hiring pipelines at once is the subject of a live federal collective action testing whether the vendor is directly liable as the employers' agent. On the litigation record, the court sustained the agent theory at the dismissal stage in 2024 and preliminarily certified a nationwide age-discrimination collective in 2025, covering applicants forty and over since September 2020, on a record in which the lead plaintiff reported more than one hundred rejections across employers using the platform. The litigation is ongoing and nothing here is an adjudicated finding of discrimination; these are allegations and procedural rulings, not a verdict.
Sources: mobleyvworkday2024
Appears on: /domains/cases/workday-screening, /pan-lab
EmpiricalA 2026 discovery ruling in the same matter held the vendor's internal bias-testing data privileged because counsel had c…
A 2026 discovery ruling in the same matter held the vendor's internal bias-testing data privileged because counsel had curated it — meaning the testing record exists and is legally unreachable, a configuration in which audit opacity is not the absence of testing but testing shielded from external verification. The case surfaces two further structural facts: a single vendor's screening model multiplied across many employer boundaries, so one learned defect can propagate as widely as the platform, and accountability diffusion between deployer and vendor, each holding part of the governance the other points to, against a survey backdrop showing assessment vendors' validation and bias-mitigation claims are often unverifiable from outside.
Sources: raghavan2020a, u2023b
Appears on: /domains/cases/workday-screening, /pan-lab
EmpiricalA graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the deployer rep…
A graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the deployer reports roughly a 90 percent reduction in time-to-hire (from about four months to about four weeks), around 50,000 candidate interview hours saved, about one million pounds in annual savings, and a 16 percent improvement in diversity. Every one of those figures is company- or vendor-reported and none is independently audited, so they are the deployer's own dashboard rather than an external measurement — which is exactly what the family's service regime looks like from inside.
Sources: bestpracticeai
Appears on: /domains/cases/unilever-hiring, /pan-lab
EmpiricalBoth vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwent a cooper…
Both vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwent a cooperative academic audit with source-code access, in which its four-fifths-rule de-biasing pipeline was found faithfully implemented — with the independence caveat that vendor staff were co-authors — and the video vendor retired its facial-analysis input under scrutiny after internal research found visual features added only about 0.25 percent predictive power, publicizing a narrow-scope external audit. The family's structural blind spot applies in full: rejected candidates never re-enter the outcome data, so the claimed quality and diversity effects are measured on hires only.
Sources: wilson2021a, maurer2021a
Appears on: /domains/cases/unilever-hiring, /pan-lab
EmpiricalA machine-learning underwriting and pricing platform using education and other alternative data operated for five years …
A machine-learning underwriting and pricing platform using education and other alternative data operated for five years under a regulator's no-action letter with a reporting obligation, and the regulator published the access results: 27 percent more applicants approved than a traditional model at 16 percent lower average APRs, with near-prime applicants (FICO 620 to 660) approved at roughly twice the rate, and gains across the tested demographic segments. This is the lending family's only regulator-verified service term. Underwriting is fully automated with no per-application human review, so the organizational levers are all upstream — model choice, the testing regime, the search for alternatives, and the reporting channel to the regulator.
Sources: consumerfinancialprotectionb2019
Appears on: /domains/cases/upstart-underwriting, /pan-lab
EmpiricalThe same deployment carries the family's most detailed public fair-lending testing record: four reports from an independ…
The same deployment carries the family's most detailed public fair-lending testing record: four reports from an independent monitorship agreed with civil-rights organizations found no close protected-class proxies quantitatively, but identified approval disparities for Black applicants, flagged a likely viable less-discriminatory alternative model, and ended in a documented methodological impasse over how hard the law requires an organization to search for such an alternative. Independently of the disparity question, adverse-action notices must give specific, accurate principal reasons for a denial regardless of the model's complexity — a governed explanation duty a complex model does not discharge by being accurate.
Sources: relmancolfaxpllc2021, consumerfinancialprotectionb2022b
Appears on: /domains/cases/upstart-underwriting, /pan-lab
EmpiricalA bank's automated credit-decisioning for a widely used consumer card was investigated by a state regulator after viral …
A bank's automated credit-decisioning for a widely used consumer card was investigated by a state regulator after viral allegations of gender bias in credit-line assignment. The regulator analyzed roughly 400,000 in-state applicants and found no unlawful discrimination on a prohibited basis — the model was cleared on the numbers. But the same investigation documented failures of explanation, customer service, and perceived transparency: applicants could not learn why they received the terms they did, front-line staff could not explain the decisions, and the resulting opacity destroyed consumer trust even though the underwriting itself was found lawful. This is the domain's cleared-but-faulted case: a statistically clean model paired with a failed duty to explain.
Sources: newyorkstatedepartmentoffina2021
Appears on: /domains/cases/goldman-apple-card, /pan-lab
EmpiricalThe lesson the cleared-but-faulted outcome carries is that a lawful, statistically clean model does not discharge the se…
The lesson the cleared-but-faulted outcome carries is that a lawful, statistically clean model does not discharge the separate duty to explain a decision. Regulators have made explicit that adverse-action notices must give specific, accurate principal reasons regardless of how complex the model is, and that a model being a black box is not a defense — checking the nearest sample-form box does not comply. The explanation and customer-service channel is therefore a distinct, separately-resourced surface that can fail on its own: an organization can pass its fair-lending testing and still fail the people it decides on by being unable to tell them why.
Sources: consumerfinancialprotectionb2022b
Appears on: /domains/cases/goldman-apple-card, /pan-lab
EmpiricalA state attorney general reached a $2.5 million settlement with a student-loan lender over its AI underwriting. The docu…
A state attorney general reached a $2.5 million settlement with a student-loan lender over its AI underwriting. The documented conduct is the domain's cleanest failure-then-mandated-governance arc: the model used a cohort-default-rate feature — a school's aggregate default rate priced into an individual applicant's terms — that disparately impacted Black and Hispanic applicants, and an immigration-status rule that automatically denied certain non-citizen applicants, while the organization ran no disparate-impact testing and gave inadequate adverse-action notices. The remedy did not fine-and-close: it mandated the missing program — model governance, disparate-impact testing, documentation, and reporting controls — so the enforcement action wrote the governance the deployment had never built.
Sources: officeofthemassachusettsatto2025
Appears on: /domains/cases/earnest-ai-underwriting, /pan-lab
EmpiricalThe mechanism the case turns on is the facially-neutral aggregate feature: a cohort default rate is a property of a scho…
The mechanism the case turns on is the facially-neutral aggregate feature: a cohort default rate is a property of a school, not of the applicant, and no input names a protected class — yet pricing a group's aggregate history into an individual's terms can carry protected-class impact, which is exactly what disparate-impact testing exists to catch. Here that testing was not done, so the impact went unmeasured until an enforcement action found it. The remedy installed the program the deployment lacked, which is the governable reading: an aggregate feature can look neutral input-by-input and still produce a disparity only outcome testing would reveal, and the absence of that testing is itself the failure.
Sources: officeofthemassachusettsatto2025, consumerfinancialprotectionb2022b
Appears on: /domains/cases/earnest-ai-underwriting, /pan-lab
EmpiricalThe strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized rollout o…
The strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized rollout of a generative-AI assistant to roughly 5,000 customer-support agents at a large software firm. Measured against a control group, the copilot raised issues resolved per hour by about 15 percent on average, and it also improved customer sentiment and agent retention. The gain, however, was sharply uneven: novice and low-skill agents improved by roughly 30 to 34 percent, agents with two months of experience performed like agents with six months and no AI, and the most experienced agents gained close to nothing, with some evidence of slight quality degradation. This is the contact-centre domain's cleanest measured benefit, and it is a distribution rather than a single number.
Sources: brynjolfsson2025a, brynjolfsson2023
Appears on: /domains/cases/fortune500-agent-copilot, /pan-lab
EmpiricalThe lesson the randomized evidence carries is skill compression: an agent-assist copilot mostly raises the floor. Becaus…
The lesson the randomized evidence carries is skill compression: an agent-assist copilot mostly raises the floor. Because almost the entire measured gain accrues to less-experienced agents and the most experienced gain close to nothing, an average productivity number overstates the effect for the agents who least need it and hides that the tool does little for the experienced while possibly costing a small amount of quality there. The governable reading is that the benefit must be measured as a distribution across agent skill, not reported as a scalar — a copilot that helps novices a great deal and experts not at all is a real and specific benefit, and describing it with one average misstates who it helps.
Sources: brynjolfsson2025a
Appears on: /domains/cases/fortune500-agent-copilot, /pan-lab
EmpiricalAn airline's customer-facing website chatbot told a customer they could claim a bereavement fare retroactively — a polic…
An airline's customer-facing website chatbot told a customer they could claim a bereavement fare retroactively — a policy that did not exist. The customer relied on the chatbot's statement, bought a ticket, and was then refused the fare by the airline's human staff. A civil-resolution tribunal found the airline liable for negligent misrepresentation and awarded damages, and in doing so rejected the airline's argument that the chatbot was a separate legal entity responsible for its own actions. The tribunal held that the organization is responsible for all the information on its website, whether it comes from a static page or a chatbot, and that a customer has no way to know which source to trust. This is the contact-centre domain's cleanest accountability ruling: the bot is a tool the company answers for, not an entity that answers for itself.
Sources: moffattv2024, sookman2024
Appears on: /domains/cases/air-canada-chatbot, /pan-lab
EmpiricalThe duty the ruling establishes is that an organization must take reasonable care that its chatbot's representations are…
The duty the ruling establishes is that an organization must take reasonable care that its chatbot's representations are accurate, because the chatbot is a tool it deploys rather than a separate entity that answers for itself. A hallucinated policy or a wrong rule stated by the bot is therefore the organization's own misrepresentation, and a posture that treats the AI as speaking only for itself does not transfer that responsibility away. The governable reading is that a customer-facing chatbot is a channel the organization is accountable for exactly as it is accountable for a page on its own website — so the accuracy control on what the bot states, and the ownership of what it says, are the organization's to build, not the bot's to carry.
Sources: sookman2024, moffattv2024
Appears on: /domains/cases/air-canada-chatbot, /pan-lab
EmpiricalAn organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds…
An organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds of customer-service chats (some 2.3 million conversations), was described as doing the equivalent work of about 700 full-time agents, cut average resolution time from about 11 minutes to under 2, was said to match human customer satisfaction, and was projected to improve profit by tens of millions. Every one of those figures was self-reported and not independently audited. Roughly a year later the same organization reversed course on quality grounds — its chief executive said cost had become too predominant an evaluation factor and the result was lower quality — and committed to always keeping a human available to customers who want one. This is the contact-centre domain's cleanest benefit-then-cost arc: the deflection numbers and the walk-back come from the same deployment.
Sources: klarnabankab2024, ivanova2025
Appears on: /domains/cases/klarna-ai-assistant, /pan-lab
EmpiricalThe lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection number reports…
The lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection number reports how many contacts the AI handled, not whether it handled them well, and a figure that is impressive on cost can hide a quality cost that only shows up later — which is what the organization's own reversal described. The survey backdrop sharpens it: most customers say they would rather not meet AI in service and fear it makes reaching a human harder, and industry analysts expect a large share of organizations to abandon plans to reduce their customer-service workforce with AI. The governable reading is to measure resolution and repeat contact against deflection rather than counting deflection as a win by itself, and to protect the path to a human as the safety valve a deflection-maximizing design tends to erode.
Sources: ivanova2025, gartner2025b
Appears on: /domains/cases/klarna-ai-assistant, /pan-lab
EmpiricalA video platform ran an unintended natural experiment on automated content moderation. When the pandemic sent its human …
A video platform ran an unintended natural experiment on automated content moderation. When the pandemic sent its human reviewers home, the platform said it would rely more on automated removal and deliberately chose over-enforcement rather than let harmful content stay up. The result, from the platform's own transparency reporting, was that automated removals more than doubled in a single quarter (to about 11.4 million videos), appeals roughly doubled, and the reinstatement rate on appeal jumped from about 25 percent to about 50 percent. The platform also withheld strikes where no human had reviewed the removal, treating the automated decision as provisional. The doubling of the reinstatement rate is the finding: it is direct evidence that the automation was making roughly twice the rate of catchable errors, and that the human review and appeals path was the loop catching them.
Sources: youtubegoogle2020
Appears on: /domains/cases/youtube-covid-enforcement, /pan-lab
EmpiricalThe lesson the natural experiment carries is that the human review and appeals path is the error-correction loop for aut…
The lesson the natural experiment carries is that the human review and appeals path is the error-correction loop for automated enforcement, not an optional add-on. Automated moderation makes errors at scale, and a doubling of the reinstatement rate when human review thinned is a measurement of those errors — they were always being made at that rate, and were visible only because the appeals queue surfaced them. Two things follow. Over-enforcement versus under-enforcement is a chosen trade-off: with review capacity cut, the organization decided which error to make, and that was a governance decision. And proactive removal acts before anyone sees the content, so an over-broad takedown is invisible unless appealed — and some removals are irreversible, as when automated systems destroyed documentation of war crimes with archival access declined, leaving no correction loop at all.
Sources: youtubegoogle2020, humanrightswatch2020
Appears on: /domains/cases/youtube-covid-enforcement, /pan-lab
EmpiricalA platform enforces its content standards with automated classifiers at a scale no human team could match, backed by a l…
A platform enforces its content standards with automated classifiers at a scale no human team could match, backed by a layered correction structure: an internal appeals process, and above it an external oversight board that issues binding decisions on the individual cases it takes and non-binding policy recommendations to the platform. In one year the board overturned the platform's original decision in around 90 percent of the cases it decided, and the platform reported implementing, in progress on, or already aligned with the large majority of the board's cumulative recommendations. This is the moderation domain's most built-out, institutionalized correction structure — layered appeals rising to an independent-adjacent external body that publishes its reasons.
Sources: metaplatforms, oversightboard2024
Appears on: /domains/cases/meta-content-enforcement, /pan-lab
EmpiricalThe reach of the correction structure is the governable limit. The roughly 90 percent overturn rate is measured on selec…
The reach of the correction structure is the governable limit. The roughly 90 percent overturn rate is measured on selected cases — the board chooses emblematic disputes to set precedent, so the figure is evidence that escalated decisions were often wrong, not a random error rate, and the overwhelming majority of automated enforcement decisions never reach the board at all. The board is funded through a platform-established trust, which makes it independent-adjacent rather than fully independent, and its policy recommendations are non-binding. The honest reading is that this correction structure is real and genuinely better than most, and its reach is bounded to the tiny fraction of cases that escalate — so the governing question is whether the correction reaches the scale of the enforcement it is meant to check.
Sources: oversightboard2024, oversightboard2025
Appears on: /domains/cases/meta-content-enforcement, /pan-lab
EmpiricalA media outlet published AI-drafted finance explainers under a human-sounding staff byline without disclosing to readers…
A media outlet published AI-drafted finance explainers under a human-sounding staff byline without disclosing to readers that the articles were machine-written. When the practice came to light, the outlet's own audit found it had to issue corrections on a majority of the AI-written articles — on the order of 41 of 77. A byline implies a human review that the reader trusts, and a correction rate that high is a direct measurement that the review the byline implied was not actually performed before publication. A later and sharper case saw another outlet publish articles under entirely fabricated author personas presented as real people, so the failure ran from undisclosed AI drafting to invented human bylines.
Sources: bonifacic2023, harrisondupre2023
Appears on: /domains/cases/cnet-ai-drafting, /pan-lab
EmpiricalEditorial AI moves the failure from a takedown to a publication, but the governable structure is the same as in moderati…
Editorial AI moves the failure from a takedown to a publication, but the governable structure is the same as in moderation: the byline is the accountability object, and it stands for a review that either happened or did not. Two things are owed to the reader — disclosure that AI was involved, and an editorial check that actually took place — and this deployment gave neither, publishing under a staff byline that implied both. When a large share of AI-drafted articles needs correction, the review was not performed, and the byline misrepresented who did the work. The governable reading is that a human byline on machine-drafted content is a claim about review and disclosure, and a high correction rate is the evidence that the claim was false.
Sources: bonifacic2023
Appears on: /domains/cases/cnet-ai-drafting, /pan-lab
EmpiricalA large automaker deployed in-line AI inspection at production scale: camera and acoustic systems that detect defects du…
A large automaker deployed in-line AI inspection at production scale: camera and acoustic systems that detect defects during assembly and feed real-time flags to the line worker via a smart device, on a line running on the order of a thousand-plus vehicles a day at a takt of under a minute per station. The system has been established as a company standard and is being extended to suppliers. The governing design is that the AI flags and a human on the line responds — the inspection is wired into a resourced response loop, including the ability to stop the line, so the benefit runs through the human response the flag triggers rather than through the model alone. The documented facts here are the system's function, the worker-interaction model, the scale, and the standardization; the deployment's benefit is reported through corporate and trade channels, and defect-rate deltas from a primary source are not public.
Sources: bmwgrouppressclub2025, metrologyandqualitynews2026, leanenterpriseinstitute
Appears on: /domains/cases/bmw-aiqx-inspection, /pan-lab
EmpiricalThe lesson the deployment carries is that an in-line inspection AI is only as good as the human-response loop it trigger…
The lesson the deployment carries is that an in-line inspection AI is only as good as the human-response loop it triggers, and that loop is the governable object. When the AI flags a defect, a resourced response — a worker with the time to check the flag and the authority to stop the line — is what turns a detection into a caught defect; without it, the flag is just a decision no one acts on. This is why the failure modes in this domain are matters of the loop's calibration rather than the model's raw accuracy: too many false alarms and operators stop responding, too much trust and they stop checking. The honest boundary is that the benefit is reported through corporate and trade channels, and no named manufacturer has publicly attributed a shipped-defect escape to its AI inspection, so the response loop is drawn as the resourced strength and its calibration as the thing to govern, not as a claim about defects that did or did not ship.
Sources: bmwgrouppressclub2025, leanenterpriseinstitute
Appears on: /domains/cases/bmw-aiqx-inspection, /pan-lab
EmpiricalA peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false alarms — …
A peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false alarms — on the order of 90 percent — through a closed operator-feedback loop: the maintenance crews investigated the alerts, labeled which were real, and the model retrained on those labels, so the false-alarm rate fell sharply over successive rounds. This is the industrial-QA domain's best-measured quantitative benefit, and it comes from an anonymized study site rather than a named-manufacturer press release, which is the pattern across this domain — the peer-reviewed magnitudes are at anonymized or smaller sites, while the named deployments report their benefit through corporate and trade channels.
Sources: hermansa2021a
Appears on: /domains/cases/heavy-industry-pdm, /pan-lab
EmpiricalThe lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that can break …
The lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that can break it, because the loop depends on the crews continuing to engage with the alerts — investigating them, labeling them, responding — and that engagement fails in two opposite directions. Alert fatigue: if too many false alarms arrive before the loop has tuned them down, crews stop trusting the alerts and stop responding, so the feedback the model needs to improve never arrives and the loop stalls. Automation bias: if crews defer to the alerts and stop applying their own judgment, the labels the model retrains on become an echo of its own calls rather than an independent check. Either way the loop degrades, so the measured benefit is contingent on the loop staying calibrated — enough trust that crews respond, enough independence that their labels still carry real judgment.
Sources: romeo2025a, wittbold2026
Appears on: /domains/cases/heavy-industry-pdm, /pan-lab
EmpiricalMachine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects t…
Machine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects that manual inspection or fixed-rule cameras would otherwise judge. In this safety-critical, regulated manufacturing setting the error trade-off is asymmetric and deliberate: a false accept — a missed defect in an injectable that reaches a patient — is a patient-safety failure, while a false reject — scrapping a good vial — is a cost, so the system is tuned to over-reject rather than risk a miss. Because the inspection sits inside a validated pharmaceutical quality process, the AI cannot simply be switched on; it must be qualified within that process, and a regulator is actively developing the framework for how AI in drug manufacturing should be validated and monitored.
Sources: veillon2023a, usfda2023
Appears on: /domains/cases/pharma-avi-inspection, /pan-lab
EmpiricalTwo governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a validated…
Two governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a validated quality process is not simply deployed but must be qualified and monitored for drift, and because the regulator's AI-specific framework is still developing, the qualification of the model's behavior over time is an emerging, not-yet-settled check rather than a solved one. Second, the human backstop: the manual inspector is what catches the false accepts the over-reject tuning is meant to avoid, so if inspectors come to defer to the AI and stop scrutinizing, that backstop erodes exactly where it matters most — the missed defect the asymmetric tuning was designed to prevent. The governable reading is that the over-reject tuning lowers the visible risk without removing it, and the qualification and the human backstop are what keep the residual risk covered.
Sources: veillon2023a, usfda2023
Appears on: /domains/cases/pharma-avi-inspection, /pan-lab
EmpiricalA state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's ri…
A state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's risk of not graduating on time and delivered the label to school staff through dashboards for about a decade. An independent, decade-scale audit found the system was wrong roughly 74 percent of the time when it predicted a student would not graduate, produced higher false-alarm rates for Black and Hispanic students, and that the deployer's own internal equity research had gone unpublished — while a survey of districts found administrators reporting no training on how to interpret a 'high risk' label. The state stopped publishing the dashboards in 2023 and said it was evaluating the system's future. The deployment is the education domain's clearest case of a risk label whose error and group disparity entered how students were seen rather than the help they received.
Sources: feathers2023, knowles2015a, wisconsindepartmentofpublici2023
Appears on: /domains/cases/wisconsin-dews, /pan-lab
EmpiricalThe lesson the case carries is that a risk label is only as good as the intervention it triggers and the training of the…
The lesson the case carries is that a risk label is only as good as the intervention it triggers and the training of the human who reads it. A label that is wrong most of the time, delivered to staff with no guidance on interpreting it, imports the model's error and its group disparity into how students are perceived rather than into a resourced response — the flag becomes a lens on the student rather than a trigger for help. Set against this, a large district's transparent, low-tech on-track indicator, built on interpretable research and paired with real intervention, accompanied a rise in graduation to a record level. The contrast locates the benefit in the intervention the indicator makes legible enough for staff to act on well, not in the sophistication of the prediction — an interpretable indicator that drives help can outperform an opaque model that only labels.
Sources: allensworth2007, feathers2023
Appears on: /domains/cases/wisconsin-dews, /pan-lab
EmpiricalA public university required students to pan their webcam around their home before an online exam, using remote-proctori…
A public university required students to pan their webcam around their home before an online exam, using remote-proctoring software that flags suspected cheating from the video. A federal court held that the pre-exam room scan was an unreasonable search under the Fourth Amendment — a first-of-its-kind ruling that a routine proctoring practice violated a student's constitutional rights in their own home. Separately, peer-reviewed measurement of automated proctoring found the software produced more face-detection failures, more red flags, and higher priority scores for darker-skinned and Black students, with no corresponding difference in actual cheating. The deployment is the education domain's clearest case of surveillance-based integrity AI whose costs — a rights violation and a demographic burden of suspicion — are each independently established.
Sources: ogletreev2022a, yoderhimes2022a
Appears on: /domains/cases/cleveland-state-proctoring, /pan-lab
EmpiricalThe lesson the case carries is that surveillance-based integrity AI is not a free default: it carries a rights cost that…
The lesson the case carries is that surveillance-based integrity AI is not a free default: it carries a rights cost that can be independently adjudicated and a demographic burden that can be measured, and both are owed a reckoning before the surveillance is imposed, not after a court or an audit finds the harm. A room scan of a student's home was held to be an unreasonable search, so the surveillance has a rights dimension a court can rule on regardless of the integrity goal. And because the software flags darker-skinned and Black students more often with no more actual cheating, and a flag is an accusation the student must answer, a disparate flag rate is a disparate burden of suspicion. The governable surfaces are the proportionality of the surveillance to the integrity problem it is trying to solve, and the measured flag rate by group.
Sources: yoderhimes2022a, ogletreev2022a
Appears on: /domains/cases/cleveland-state-proctoring, /pan-lab
EmpiricalA parcel carrier's route-optimization system is a documented operations-research success: it re-optimizes delivery route…
A parcel carrier's route-optimization system is a documented operations-research success: it re-optimizes delivery routes across the fleet and was reported to save on the order of 100 million miles and about 10 million gallons of fuel a year, a genuine and peer-reviewed efficiency gain. The same system that computes the efficient route also dictates it to the driver and monitors adherence through vehicle telematics, so the efficiency is enforced through workplace surveillance — the optimization and the monitoring are one system, and the driver's discretion over how to run the route is what it replaces. The benefit is real and measured in miles and fuel; the cost is the driver autonomy the enforcement removes and the surveillance the enforcement requires.
Sources: holland2017, levy2023
Appears on: /domains/cases/ups-orion-routing, /pan-lab
EmpiricalThe lesson the case carries is that an optimization which manages the worker executing it couples the efficiency gain to…
The lesson the case carries is that an optimization which manages the worker executing it couples the efficiency gain to a cost the efficiency metric does not see: the worker's autonomy, and the surveillance required to enforce the plan. The system measures miles and fuel, not whether the pace it sets is feasible for a person or whether the monitoring it requires is proportionate — so the governable surfaces are whether the optimization internalizes the human executing it, meaning a route that is feasible and humane rather than merely optimal on paper, and whether the surveillance that enforces it is governed rather than treated as a free byproduct of routing. An optimization is a success on its own terms and can still externalize a cost onto the worker that never appears in the miles-and-fuel number it reports.
Sources: levy2023, holland2017
Appears on: /domains/cases/ups-orion-routing, /pan-lab
EmpiricalA warehouse operation's algorithmic management pairs a genuine, peer-reviewed human-robot picking benefit — robots and w…
A warehouse operation's algorithmic management pairs a genuine, peer-reviewed human-robot picking benefit — robots and workers collaborating to raise throughput, documented in the operations-research literature — with a documented injury-productivity trade-off. When the algorithm sets the pace of the physical work, a federal safety regulator cited the operation for exposing workers to ergonomic hazards, and a legislative inquiry tied the speed the system demands to warehouses it described as uniquely dangerous. The productivity gain and the worker-injury risk are therefore coupled: the same pace that raises units per hour is the pace regulators and the inquiry connected to injury. The benefit is real and the injury cost is separately documented, one in the OR literature and one in safety-inspection findings and a legislative report.
Sources: allgor2023, ussenatecommitteeonhealth2024, usdepartmentoflabor2023a
Appears on: /domains/cases/amazon-fulfillment-management, /pan-lab
EmpiricalThe lesson the case carries is that when an algorithm sets the pace of physical work, the productivity metric it optimiz…
The lesson the case carries is that when an algorithm sets the pace of physical work, the productivity metric it optimizes — units per hour — cannot see the cost the pace imposes on the body executing it. The injury shows up in safety-inspection data and a legislative inquiry, not on the throughput dashboard, so a productivity number can rise while the cost accumulates unrecorded on the metric that reports success. The governable question is whether the pace-setting internalizes the worker's safety, treating a sustainable rate as part of what 'optimal' means, or externalizes it as an injury the metric never records. Ethnographic research describes this algorithmic management as a 'game' whose rules the worker cannot change, which is what makes the pace a management decision the organization owns rather than a fact of the work.
Sources: cheon2025, ussenatecommitteeonhealth2024
Appears on: /domains/cases/amazon-fulfillment-management, /pan-lab
EmpiricalA federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short …
A federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short speech sample, as one input into the credibility assessment of their claimed origin. The tool's reliability is limited: government-reported recognition is around 80 percent for Arabic — roughly a 20 percent error rate — and computational linguists judge separating some closely related language varieties close to hopeless. The agency's own caseworkers describe the tool as only a rough compass, too imprecise to resolve the hard cases, and its outputs as clues rather than determinations. Used honestly as one clue among several it is defensible; the documented risk is that an imprecise output acquires more authority than its accuracy supports, in a determination where the state's tool is set against the applicant's own account of who they are.
Sources: lulamae2022a, scheel2024a
Appears on: /domains/cases/bamf-dias-dialect, /pan-lab
EmpiricalTwo governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determination. First…
Two governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determination. First, whether the tool's documented imprecision actually bounds the weight it carries: a rough compass treated as one is honest, but the same output can harden into a credibility finding it cannot support once a phrase like the software indicates a particular origin enters the record and confronts the applicant. Second, whether the applicant can see and contest the signal: in asylum determinations the person with the most at stake and the most knowledge of their own origin is often unable to see or challenge the AI's estimate, so the correction that would catch an error is severed on exactly the side that holds the truth. The governable reading is that reliability must bound authority, and the affected person must be able to contest a signal used against them.
Sources: scheel2024a, lulamae2022a
Appears on: /domains/cases/bamf-dias-dialect, /pan-lab
EmpiricalA government's immigration-enforcement triage algorithm identifies and recommends people for enforcement actions — retur…
A government's immigration-enforcement triage algorithm identifies and recommends people for enforcement actions — returns, bail conditions, casework — drawing on sensitive data including detention, health, vulnerability, and location-monitoring records. Uncovered through roughly a year of freedom-of-information litigation, its training materials show an asymmetric override design: officials must record a justification for rejecting a recommendation but not for accepting one. That design builds a rubber-stamping incentive into the workflow — accepting the algorithm is frictionless, overriding it requires work — so the human in the loop is nominal rather than a real check. It is the corpus's clearest documented instance of automation bias engineered into an agency workflow, in one of the highest-stakes enforcement settings a state operates.
Sources: privacyinternational2024b, privacyinternational2024a
Appears on: /domains/cases/home-office-ipic, /pan-lab
EmpiricalThe lesson the case carries is that nominal human oversight is not real oversight. An asymmetric override — where accept…
The lesson the case carries is that nominal human oversight is not real oversight. An asymmetric override — where accepting the algorithm's recommendation is frictionless and rejecting it requires a recorded justification — engineers automation bias into the process by making deference the path of least resistance, so the claim that a human makes the final decision can be true and empty at once. Two governable surfaces follow. Whether the review is genuinely symmetric: the official as free and as prompted to reject as to accept, so an error is as likely to be caught as waved through. And whether the affected person is told the AI is used and can contest it: applicants are frequently not told, which severs the correction on the side that could challenge the recommendation, so the one check that survives the asymmetric override — the person it is about — is cut out too.
Sources: privacyinternational2024c, privacyinternational2024a
Appears on: /domains/cases/home-office-ipic, /pan-lab
EmpiricalA large automaker developed an in-house deep-learning system to detect hairline cracks in pressed sheet-metal parts, tra…
A large automaker developed an in-house deep-learning system to detect hairline cracks in pressed sheet-metal parts, trained on several terabytes of images drawn from seven presses at its home plant plus several sister plants, in development since mid-2016 and tested for series deployment. The documented change is a generational replacement: the system takes over an inspection duty previously performed by manual visual checks plus fixed-rule camera systems, rather than augmenting a human inspector's judgment on each part. The record — a reprint of the manufacturer's own press material with its CIO quoted — documents the development lineage, the data scale, and what the system replaced; it publishes no quantitative defect-rate figures, so the deployment's benefit magnitude is a corporate claim, not an audited measurement.
Sources: justauto2018
Appears on: /domains/cases/audi-press-shop-inspection, /pan-lab
EmpiricalThe governance shape of this deployment is inheritance rather than assistance: by replacing the manual visual check and …
The governance shape of this deployment is inheritance rather than assistance: by replacing the manual visual check and the fixed-rule camera generation, the learned system inherits the whole inspection duty for the defect class it covers, so there is no per-part human judgment running alongside it to catch what it misses. Its training data is pooled across presses and plants, which means one model's blind spots are correlated across every line it inspects. The failure regime is mechanism-level — drift as dies wear and parts change, complacency over an inspection nobody re-performs — because no named manufacturer, including this one, has publicly attributed a shipped-defect escape to its AI inspection.
Sources: justauto2018
Appears on: /domains/cases/audi-press-shop-inspection, /pan-lab
EmpiricalA rail service provider operates sensor-based predictive maintenance on a high-speed fleet under a priced availability c…
A rail service provider operates sensor-based predictive maintenance on a high-speed fleet under a priced availability contract: roughly 300 sensors per train read at five-minute intervals (on the order of a million readings per train-year), overlaid with human-written failure reports, maintained against a promise that refunds the full fare if a journey is delayed more than fifteen minutes. The documented results are only one noticeably delayed journey in 2,300 (by five minutes) and discovered failure signatures such as an engine-temperature pattern preceding failure by three days. The fleet outcome is documented in a vendor-side trade case study; the analytics' own precision is not published, so the model-level figures remain unstated while the operational outcome is on the record.
Sources: rcrwirelessnews2016
Appears on: /domains/cases/siemens-renfe-velaro-pdm, /pan-lab
EmpiricalThe governance shape of this deployment is uptime-as-contract: the party that operates and tunes the analytics is the se…
The governance shape of this deployment is uptime-as-contract: the party that operates and tunes the analytics is the service provider who pays for misses under the refund promise, so the incentive to prevent a delay is priced into the same organization that holds the model levers - an alignment the domain's other deployments lack. Two documented dependencies temper it: the failure signatures were discovered by joining sensor streams to human-written failure reports, so the discovery loop runs on documentation crews write for their own purposes; and a continuous sensor stream is the analytics' only view of the machine, so a failing sensor and a failing train arrive looking the same until someone goes and looks.
Sources: rcrwirelessnews2016
Appears on: /domains/cases/siemens-renfe-velaro-pdm, /pan-lab
EmpiricalA large urban school district operationalized a transparent ninth-grade indicator - course credits earned plus no more t…
A large urban school district operationalized a transparent ninth-grade indicator - course credits earned plus no more than one core-course failure - from consortium research showing it predicts high-school graduation with about 85 percent accuracy, and wired it to school-level attention rather than to an opaque score. District graduation rates subsequently rose to record highs. The indicator is a rule anyone can read: a teacher can explain to a student exactly why they are off-track and exactly what would change it, so the contest-and-correction loop that opaque early-warning deployments sever is open by construction.
Sources: allensworth2007
Appears on: /domains/cases/cps-freshman-ontrack, /pan-lab
EmpiricalThe documented limits are as instructive as the result. The indicator's accuracy and the district's graduation rise are …
The documented limits are as instructive as the result. The indicator's accuracy and the district's graduation rise are associational at district scale - no randomized trial assigns schools to use it - and the benefit mechanism runs through the intervention, not the flag: an indicator wired to attention still depends on the attention being resourced, and the research base's central finding is that what predicted graduation was a condition schools could act on (freshman-year course performance), not a fixed trait of the student. The rule's power is that it points at something changeable, and the district's practice is what changed it.
Sources: allensworth2007
Appears on: /domains/cases/cps-freshman-ontrack, /pan-lab
EmpiricalA storied sports outlet published product-review articles under entirely fabricated author personas - invented names, AI…
A storied sports outlet published product-review articles under entirely fabricated author personas - invented names, AI-generated headshots, fictional biographies - produced by a third-party content contractor, with the AI involvement disclosed to no reader. An investigation surfaced the personas by reading the public site; the articles were then deleted rather than corrected, the outlet attributed the content to the contractor, and the parent company's chief executive was subsequently fired. The corroborated record documents the fabrication, the deletion, the contractor attribution, and the executive consequence.
Sources: harrisondupre2023, npr2023
Appears on: /domains/cases/sports-illustrated-advon, /pan-lab
EmpiricalThe governance failure ran across an organizational seam: the drafting, the bylines, and the personas were produced by a…
The governance failure ran across an organizational seam: the drafting, the bylines, and the personas were produced by a contractor, and the outlet's editorial function demonstrably did not operate across that boundary - the fabrication was discovered by outside investigation, not by any internal check, and the accountability afterward ran through contract and employment rather than through any editorial process. A byline makes two claims to the reader - that a person produced this, and that the outlet's review stands behind it; this deployment fabricated the first and vacated the second, and the deletion afterward removed the evidence rather than correcting the record.
Sources: harrisondupre2023, npr2023
Appears on: /domains/cases/sports-illustrated-advon, /pan-lab
EmpiricalA parcel firm's customer-facing support chatbot, after a system update, was prompted by a customer into swearing and int…
A parcel firm's customer-facing support chatbot, after a system update, was prompted by a customer into swearing and into composing a poem calling its own operator the worst delivery firm in the world. The firm attributed the behavior to the update and disabled the AI element immediately. The documented governance facts are exactly two: the update preceded the behavior, and the off switch worked - the firm learned of the incident from the customer's viral post rather than from any release gate, but the disablement was immediate once it knew.
Sources: itvnews2024
Appears on: /domains/cases/dpd-uk-chatbot, /pan-lab
EmpiricalThe failure shape is a change-management regression, not a wrong policy or a deflection metric: guardrails that had held…
The failure shape is a change-management regression, not a wrong policy or a deflection metric: guardrails that had held in production stopped holding after a change, publicly, within hours, in a channel that talks to anyone. What the deployment lacked was a release gate between the update and the public - the constraint layer's behavior after the change was tested by a customer with a prompt, not by the firm with a suite - and the discovery path ran through screenshots of one conversation going viral.
Sources: itvnews2024
Appears on: /domains/cases/dpd-uk-chatbot, /pan-lab
EmpiricalSince May 2018, New York City's Administration for Children's Services has scored every open child-protection investigat…
Since May 2018, New York City's Administration for Children's Services has scored every open child-protection investigation at day 10 with an in-house machine-learning model (documented as the ASAP Tool / Severe Harm Predictive Model), rank-ordering cases by predicted likelihood of substantiated physical or sexual abuse within 24 months to fill a quality-assurance review worklist of about 3,000 of roughly 50,000 investigations a year; the LL35 register states scores are not shared with staff in the QA unit or the investigative unit, families and their attorneys are not told when a case is flagged, ACS told state auditors there would be 'no basis for a complaint' about a predictive model on an individual case, and the agency's own internal audit acknowledged the training data likely included implicit and systemic biases, that geographic variables may act as partial proxies for race, and that flag predictions are more likely to be incorrect than correct.
Sources: nycofficeoftechnologyandinno2026, lecher2025, newyorkstatecomptroller2023
Appears on: /domains/cases/nyc-acs-qa-risk-algorithm
EmpiricalThe May 2026 'Access Denied' report by the NYC Department of Investigation — the Charter-mandated inspector general for …
The May 2026 'Access Denied' report by the NYC Department of Investigation — the Charter-mandated inspector general for ACS — documents that five provisions of NY Social Services Law, as applied by state OCFS, routinely deny, limit, or delay DOI's access to ACS child-welfare records, with unfounded-report and CARES records prohibited entirely; DOI was barred from the full case history in 17 of the 18 child fatalities with prior ACS involvement reported to it in 2025 (13 of 16 in 2024; 19 of 25 in 2023). The report does not mention the algorithm — the linkage is a topological inference — but in 2025 ACS expanded the model to score CARES alternative-response cases, the record class state law prohibits DOI from accessing; the NYC Council's GUARD Act (passed unanimously November 25, 2025) legislated an Office of Algorithmic Data Accountability whose implementation status remains unverified as of mid-2026.
Sources: newyorkcitydepartmentofinves2026, nycofficeoftechnologyandinno2026, statescoop2025
Appears on: /domains/cases/nyc-acs-qa-risk-algorithm
EmpiricalThe ACS severe-harm model trains on and scores exclusively from the agency's own administrative records (including hotli…
The ACS severe-harm model trains on and scores exclusively from the agency's own administrative records (including hotline call counts and durations, prior involvement, and geography), and the intervention its flags trigger — extra interviews, collateral contacts, service referrals, consults, and QA documentation follow-ups — writes new activity into those same records, enriching the prior-involvement features of any future report on the family; ACS has never studied what effect the flag-triggered extra review has on downstream case outcomes, including whether a child is ultimately removed from the home (while telling state auditors it produces quarterly internal reports tracking the model's use in the QA program), kept no logs of model performance evaluations or updates as of the 2019-2022 state-audit fieldwork, and its claim that the model outperformed experienced caseworkers with fewer false positives and more race/ethnicity equity is an agency self-claim with no published methodology or independent verification.
Sources: lecher2025, newyorkstatecomptroller2023
Appears on: /domains/cases/nyc-acs-qa-risk-algorithm
EmpiricalColorado's legislature built a distinctive oversight topology — a standing independent statutory ombudsman (created 2010…
Colorado's legislature built a distinctive oversight topology — a standing independent statutory ombudsman (created 2010; an independent judicial-department agency since 2016 under C.R.S. 19-3.3) was handed a funded, nine-criteria audit mandate (HB 24-1046, signed May 28, 2024; $109,392 appropriation) over the statewide Family Safety and Family Risk Assessment tool suite used by 64 county-administered agencies — and the resulting ICF audit (176 pages, dated February 27, 2026, released March 2, 2026, with 50 recommendations across nine directives) found the tools well-aligned with policy on paper but inconsistently implemented: only 32% of surveyed caseworkers complete assessments in real time, vignette agreement fell to 59–63% on neglect and poverty scenarios versus 91–96% on abuse scenarios, and the audit recommends redefining safety and risk with observable behavior-based criteria and revising or replacing the actuarial risk tool.
Sources: icfincorporated2026, steffen2026, coloradogeneralassembly2024a, coloradorevisedstatutes2024
Appears on: /domains/cases/colorado-safety-risk-tools-audit
EmpiricalColorado's statutory audit loop fired end to end on the way up — the Child Protection Ombudsman's complaint-stream obser…
Colorado's statutory audit loop fired end to end on the way up — the Child Protection Ombudsman's complaint-stream observations became testimony to the 2023 Child Welfare System Interim Study Committee (including its brief stating the Family Safety Assessment had never been validated since its 1999 inception), the testimony became HB 24-1046 requiring the ombudsman rather than the child welfare agency to procure a third-party audit, and ICF's competitively procured audit returned to named legislative committees by the March 1, 2026 statutory deadline with four public information sessions following — but as of July 2026 the loop has produced statute, budget, a public audit artifact, and public information only: no follow-up legislation has been enacted, no formal CDHS response has been identified, and the tools remain in statewide operation unchanged, a decade after the state auditor's October 2014 performance audit flagged overlapping assessment and documentation weaknesses without producing redesign.
Sources: coloradogeneralassembly2024, childprotectionombudsmanofco2026, steffen2026, officeofcoloradoschildprotec2023, coloradoofficeofthestateaudi2014, icfincorporated2026
Appears on: /domains/cases/colorado-safety-risk-tools-audit
EmpiricalThe ICF audit found Colorado's actuarial Family Risk Assessment 'heavily weights prior reports, which inflate risk score…
The ICF audit found Colorado's actuarial Family Risk Assessment 'heavily weights prior reports, which inflate risk scores for incidents that occurred long ago, and disproportionately impacts minority families and perpetuates systemic bias' — with historical risk factors that 'disproportionately elevate scores for families of color' and a DV history item that codes any past incident, even decades old, as recent, penalizing survivors — while the record substrate degrades the oversight signal meant to catch exactly this: 68% of surveyed caseworkers cannot complete assessments in real time (documentation is back-filled into Trails after decisions), and race and ethnicity are so inconsistently documented that the audit says the disproportionality analysis requested by the legislature was limited; these are structural and qualitative audit findings, as no administrative-data disparity analysis has been published.
Sources: icfincorporated2026
Appears on: /domains/cases/colorado-safety-risk-tools-audit
EmpiricalFamily-Match, a proprietary two-sided 'relational fit' adoption matching algorithm built for Adoption-Share by former eh…
Family-Match, a proprietary two-sided 'relational fit' adoption matching algorithm built for Adoption-Share by former eharmony researchers, produced 1 known adoption in Virginia's two-year test (an official's statement to the AP; VDSS said the tool 'had not proven effective'; the pilot's end is undated in the record) and 2 adoptions in Georgia's year-long pilot ended October 2022; caseworkers in Florida, Georgia, and Virginia said it often led them to unwilling families, and Virginia social workers were perplexed that the algorithm seemed to match all the children with the same group of parents. In Florida the vendor's own quarterly report claimed 603 placements yielding 431 adoptions over five years — figures partner agencies could not verify: FamiliesFirst Network's records showed 76 Family-Match placements with no documented adoption plus 3 failed trial placements since 2019, and Children's Network of Southwest Florida counted 22 matches and 8 adoptions in five years while making hundreds of matches and hundreds of adoptions without the tool over the same period.
Sources: hoandburke2023a, fortuneassociatedpressrepubl2023
Appears on: /domains/cases/family-match-adoption-share
EmpiricalThe outcome evidence for Family-Match lived in a vendor-owned, credit-claiming data store: Virginia officials said that …
The outcome evidence for Family-Match lived in a vendor-owned, credit-claiming data store: Virginia officials said that once families' data was entered 'Adoption Share owned the data,' assistant director Traci Jones said 'We did not have access to the algorithm even after it was requested,' and agencies 'couldn't explain Family-Match's self-reported data.' The vendor's April 2023 'confidential' user guide instructed caseworkers not to delete cases matched outside the tool but to document them in the system 'so that Adoption-Share could refine its algorithm and follow up with the families'; Miami's Citrus Family Care Network said vendor staff asked social workers to have parents register in the tool even when it played no role in the adoption, and Georgia's spokesperson said Family-Match could claim credit for pairings already in its system. In Georgia the vendor-owned store holds whether foster youth have been sexually abused, the gender of their abuser, criminal records, and whether they identify as LGBTQIA — data typically restricted to secured child protective services case files — and two Florida agencies fed the system data with no contract and would not say how children's data was secured.
Sources: hoandburke2023a, fortuneassociatedpressrepubl2023
Appears on: /domains/cases/family-match-adoption-share
EmpiricalFamily-Match's four-state lifecycle turned on procurement authority actions made while states were structurally dependen…
Family-Match's four-state lifecycle turned on procurement authority actions made while states were structurally dependent on the vendor's own ledger for outcome evidence: Georgia ended its pilot on null results in October 2022, then — after Ramirez met with the governor's office and lobbied a statehouse committee — signed a new agreement in July 2023 for free; Virginia dropped the matching pilot as 'not proven effective' and by 2022 awarded Adoption-Share a larger contract for the Faster Families Highway recruitment portal ($188,000 budgeted / $212,546 spent SFY2023, $246,100 planned SFY2024, all 120 local departments enrolled by December 31, 2022, with policies and procedures — including family removal and response-time expectations — drafted only from February 2023, after statewide enrollment); Florida converted a philanthropy-funded rollout to a $350,000 DCF contract in October 2023 and added a Florida DOH contract for a medically-complex-children algorithm; and Tennessee, the only state whose reviewers formally questioned pre-deployment why the tool needed certain sensitive data points and how they influenced the match score, scrapped the rollout before it began. No state has ever published an evaluation of the matching pilot.
Sources: hoandburke2023a, virginiadepartmentofsocialse2023
Appears on: /domains/cases/family-match-adoption-share
EmpiricalFive US states (Michigan 2001, Minnesota 2001/2006, Maryland 2009, Texas 2013, Missouri 2021) operate 'Birth Match' prog…
Five US states (Michigan 2001, Minnesota 2001/2006, Maryland 2009, Texas 2013, Missouri 2021) operate 'Birth Match' programs in which an identity record-join between new birth registrations and a registry of parents with prior terminations of parental rights, serious-harm findings, or specified convictions fires a mandatory child-protective response with no risk score, no threshold, and no screening discretion: Michigan's current policy (PSM 712-2, reissued 2026-04-01) permits screen-out only for an inaccurate identity match or an already-open case and otherwise requires the referral to be screened in and assigned for investigation, and Maryland's 2025 response policy limits screen-out to a non-qualifying conviction or an adopted child while requiring a 24-hour intake coded 'Risk of Harm: birth match' — a mandatory response with no screening discretion, formally structured in Maryland as a voluntary non-CPS assessment families may refuse, with tracing and escalation duties attached. The only tunable parameters anywhere are registry scope and lookback: 2 years in Texas, 10 in Maryland and Missouri, unlimited in Michigan (records to 1978) and Minnesota.
Sources: cohen2022, michigandepartmentofhealthan2026, marylanddepartmentofhumanser2025, minnesotarevisorofstatutes2021
Appears on: /domains/cases/us-birth-match
EmpiricalThe Birth Match trigger is self-arming: the registry that fires it is populated by the system's own outputs, so a birth-…
The Birth Match trigger is self-arming: the registry that fires it is populated by the system's own outputs, so a birth-match response that ends in a new termination of parental rights writes the parent back onto the match list — permanently in Michigan and Minnesota, which have no lookback limit or retirement mechanism — and every subsequent birth re-fires with strictly increasing history; Michigan additionally adds manual list entries for severe cases without any termination. Maryland law accelerates the loop by letting the agency skip reasonable reunification efforts on account of a prior involuntary termination, and the 2025-26 Maryland legislative fight (HB 944 of 2025, died without a committee vote; HB 48 of 2026, heard 2026-01-29 and dead in committee without a vote when the session adjourned sine die on April 13, 2026) targeted that waiver, not birth-match repeal — with Civil Rights Corps testifying that the combination is 'functionally state sterilization' and coerces parents into 'voluntary' relinquishment, an advocacy characterization in the legislative record.
Sources: cohen2022, richardsonandsteckelcivilrig2026, marylandgeneralassembly2026, michigandepartmentofhealthan2025
Appears on: /domains/cases/us-birth-match
EmpiricalThe Birth Match automation runs without ownership or evaluation: match volumes fell by more than half in Michigan (1,186…
The Birth Match automation runs without ownership or evaluation: match volumes fell by more than half in Michigan (1,186 in FY2019 to 515 in FY2021) and Maryland (243 in CY2019 to 124 in CY2020) without the operating agencies noticing or being able to explain the change; fewer than 10 percent of matched families in any state with data received services (Maryland: 5 of 124 matched families in the source's mixed FFY2020/CY2020 pairing); Maryland's unanimous 2018 expansion statute ordered an independent evaluation of the match's sensitivity, specificity, and predictive value that a public-records request found was never implemented; and the only oversight that demonstrably changed the system — child fatality review teams auditing the matcher's misses — widened the trigger both documented times. Nearly all of these figures flow through one proponent researcher's ad hoc public-records requests (AEI 2022), which the author herself flags as limited and of dubious accuracy; no state publishes birth-match data in any regular report.
Sources: cohen2022, cohen2022a, marylandgeneralassembly2018, baltimorecitychildfatalityre2017
Appears on: /domains/cases/us-birth-match
EmpiricalDC's Child and Family Services Agency published a 16-page pre-deployment AI Values Alignment Report (one of only three d…
DC's Child and Family Services Agency published a 16-page pre-deployment AI Values Alignment Report (one of only three district-wide under Mayor's Order 2024-028) for CORA, a staff-facing policy chatbot launched June 16, 2025 inside its STAAND case-management system, whose centerpiece guardrail is a read-side exclusion — 'the chatbot does not draw from confidential case files' and has no access to STAAND case data; by December 2025 the agency's own tip sheets documented a CORA phone app that ingests case documents, photos of handwritten notes, and voice dictation and saves AI-drafted contact notes into the STAAND case record after worker approval — the write direction the no-case-files rule never governed, superseding the May 2025 report's dated statement that generated text is 'guidance, not text to be used by the employee as part of case documentation.'
Sources: dcchildandfamilyservicesagen2025b, dcchildandfamilyservicesagen2025, dcofficeofthechieftechnology2026
Appears on: /domains/cases/dc-cfsa-cora-chatbot
EmpiricalCFSA's May 2025 AI Values Alignment Report concedes that per-output human validation of CORA is impossible — asked wheth…
CFSA's May 2025 AI Values Alignment Report concedes that per-output human validation of CORA is impossible — asked whether humans can review and approve AI outputs before they are enacted, it answers 'No,' with validation occurring before content enters the knowledge source (SME validation of every document, annual committee re-review of the corpus) and afterward via quality assurance (monthly internal accuracy reports and a real-time answer-flagging channel, none published); by December 2025 the tool's own landing card carried the in-product warning 'The Ask Feature is under construction and answers may not be fully accurate. Consult your supervisor as needed,' delegating per-query vigilance to workers and supervisors six months after launch.
Sources: dcchildandfamilyservicesagen2025b, dcchildandfamilyservicesagen2025a
Appears on: /domains/cases/dc-cfsa-cora-chatbot
EmpiricalCORA operates as a single-agency closed loop with no external examination of the running system: its knowledge authors, …
CORA operates as a single-agency closed loop with no external examination of the running system: its knowledge authors, tool owners, operators, and oversight committees are all CFSA units; no OIG or auditor review, no published accuracy or usage data, and no DC Council finding exists as of mid-2026; the deployment sits in CFSA's first era without a court monitor in three decades (LaShawn A. oversight ended 2021, final closure 2022); the 25,000 queries per month figure is a budget ceiling, not observed usage; and every quantitative outcome figure (45 minutes saved per intake report, one to four hours per case, three-week feature delivery at a claimed 20x lower cost) is an unaudited vendor claim from a customer story that never names CORA, so CORA-specific claims rest on agency sources alone.
Sources: dcchildandfamilyservicesagen2025b, microsoft2025, dcchildandfamilyservicesagen2021
Appears on: /domains/cases/dc-cfsa-cora-chatbot
EmpiricalIn Richmond, Virginia, Predict-Align-Prevent's open-source place-based model (built with Urban Spatial; report byline Ke…
In Richmond, Virginia, Predict-Align-Prevent's open-source place-based model (built with Urban Spatial; report byline Ken Steif, Matthew D. Harris, and Sydney Goldstein, 2019) reported its highest risk tier capturing about 70% of held-out child-maltreatment events against about 35% for a kernel-density baseline, while covering about 10% of city land area holding roughly 48,500 residents including about 8,200 children — builder-generated figures from a commissioned report — and the builders' own published fairness audit stated verbatim that the meta-model 'generalizes well across neighborhoods of varying poverty rates, but does not generalize well across neighborhoods of varying race,' despite race and income being excluded from the feature set.
Sources: steif2019, urbanspatial2019
Appears on: /domains/cases/predict-align-prevent-geospatial
EmpiricalNew Hampshire's 2018-2023 federal Community Collaborations cooperative agreement is the only documented operational plan…
New Hampshire's 2018-2023 federal Community Collaborations cooperative agreement is the only documented operational planner use of Predict-Align-Prevent's maps: state narratives confirm the mapping (funded by Casey Family Programs, using address-level inputs accessed through stewarded databases) was completed for all three service areas before sites received implementation resources, and the March 2024 ACF/OPRE grantee profile documents that Community Implementation Teams used PAP-identified areas of need to target family outreach and issued funded requests for proposals with stipends — documented resource-routing behavior change, with no independent outcome evaluation of maltreatment effects in any deployment.
Sources: administrationforchildrenand2024, stateofnewhampshire2020
Appears on: /domains/cases/predict-align-prevent-geospatial
EmpiricalThe 2016 Fort Worth study (Daley et al., Child Abuse & Neglect), trained on 2013 state child-welfare substantiations and…
The 2016 Fort Worth study (Daley et al., Child Abuse & Neglect), trained on 2013 state child-welfare substantiations and Fort Worth police data, reported the top 10% of grid cells capturing 52% of 2014 maltreatment cases against 43% for a conventional hotspot model — figures independently corroborated verbatim by the Marchment and Gill (2021) Crime Science systematic review, which also found it to be the single child-maltreatment application of risk terrain modeling in the reviewed literature; Fort Worth remained a published study plus local advocacy, with no evidence of operational resource allocation by the maps.
Sources: daley2016, marchmentandgill2021
Appears on: /domains/cases/predict-align-prevent-geospatial
EmpiricalIn December 2025 the Massachusetts Department of Transitional Assistance piloted an Accenture-built tool that transcribe…
In December 2025 the Massachusetts Department of Transitional Assistance piloted an Accenture-built tool that transcribes SNAP eligibility calls in real time and generates a structured, caseworker-editable summary that is saved into BEACON, the state's benefits eligibility system of record; full transcripts are not retained, only the summaries, according to technical documentation reviewed by The Shoestring, and the summarization prompt was withheld as proprietary. About 400 calls had been processed by the April 2026 reporting, against roughly 45,000 calls connected to staff per month, in a record system that fed 31,390 SNAP application dispositions in December 2025 alone. No independent evaluation, inspector-general audit, or completed privacy impact assessment of the tool is on record.
Sources: theshoestring2026, massachusettsdepartmentoftra2025
Appears on: /domains/cases/massachusetts-dta-call-summaries
EmpiricalOversight of the DTA call summarizer ran through the labor channel: SEIU Local 509, representing DTA call-center workers…
Oversight of the DTA call summarizer ran through the labor channel: SEIU Local 509, representing DTA call-center workers among roughly 9,000 state employees, reached an agreement with DTA — pre- or early-deployment; the record does not establish it preceded the December 2025 rollout — that made worker use voluntary and protected jobs, without addressing record provenance or client-side safeguards. The formal privacy apparatus sat empty: none of the nine AI use cases Massachusetts disclosed from its internal inventory of at least 40, the DTA summarizer included, reported a completed privacy impact assessment; the tool's interaction-data plan was listed as still to be defined while it was in production; and details on the other 31 use cases were withheld until the Supervisor of Records ordered them submitted for in camera review.
Sources: theshoestring2026, massachusettsexecutiveoffice2026
Appears on: /domains/cases/massachusetts-dta-call-summaries
EmpiricalThe record store the AI summaries enter was itself contested while the pilot ran: in California v. USDA, No. 3:25-cv-063…
The record store the AI summaries enter was itself contested while the pilot ran: in California v. USDA, No. 3:25-cv-06310 (N.D. Cal.), a 21-state-plus-DC coalition including Massachusetts obtained preliminary injunctions on October 15, 2025 and February 27, 2026 blocking USDA from cutting SNAP funding over states' refusal to hand over personal SNAP applicant and recipient data, the court holding the proposed data protocol would likely permit sharing beyond the entities allowed under 7 U.S.C. 2020(e)(8). Both orders are preliminary and the litigation is live; whether AI-generated call summaries held in BEACON fall within the demanded data is unresolved, and the Massachusetts AG's office declined to comment on that question.
Sources: massachusettsattorneygeneral2026, jurist2026, theshoestring2026
Appears on: /domains/cases/massachusetts-dta-call-summaries
EmpiricalIllinois DCFS rolled out a vendor natural-language-processing layer statewide over its legacy SACWIS case-management sys…
Illinois DCFS rolled out a vendor natural-language-processing layer statewide over its legacy SACWIS case-management system beginning in 2023, provisioning roughly 7,500 users with a DCFS-estimated 20 percent administrative-time saving; the agency's federal Annual Progress and Services Report states that all notes in SACWIS can be mined, rendered as per-case Risks and Strengths overviews with click-through to source notes. By the vendor CEO's November 2025 congressional testimony, more than 6,000 Illinois staff use the tool and it saves each social worker about five hours per week — vendor figures delivered in a hearing whose majority summary records no AI-risk scrutiny — and no independent accuracy evaluation, published usage metric, or audit of the deployment has been located.
Sources: elisco2025, illinoisdepartmentofchildren2024, governmenttechnology2023, housewaysandmeanscommittee2025
Appears on: /domains/cases/illinois-dcfs-augintel
EmpiricalFrom the 2023 announcement onward the same platform carried a second, management-facing channel reading the same note st…
From the 2023 announcement onward the same platform carried a second, management-facing channel reading the same note store: early-warning signs, caseworker safety, practice-model fidelity, and statewide trends per the launch release; compliance-related data collected in the background from the notes workers already document, service-outcome mining tied to funding accountability, and vendor-analytics deep dives informing where programs should be established or eliminated, per the vendor CEO's 2025 testimony; and a named Chapin Hall partnership using machine learning over case notes to measure practitioners' use of Motivational Interviewing. The associated documentation-behavior effect on the workers whose notes are mined is a structural inference, not a documented finding.
Sources: augintel2023, elisco2025, chapinhallattheuniversityofc2025
Appears on: /domains/cases/illinois-dcfs-augintel
EmpiricalIn December 2017 the same agency, Illinois DCFS, terminated its Eckerd Rapid Safety Feedback predictive-analytics pilot …
In December 2017 the same agency, Illinois DCFS, terminated its Eckerd Rapid Safety Feedback predictive-analytics pilot — a 366,000-dollar sole-source arrangement that scored more than 4,100 children at 90-percent-plus probability of death or serious injury while assigning low risk to two children who died — after a joint OEIG/DCFS-OIG report found the arrangement had been misclassified as a grant rather than a no-bid contract. That shutdown is the only oversight event at this agency documented to have changed system behavior, and it preceded the current deployment's deliberately score-free extraction design.
Sources: theimprint2017, governmenttechnologya
Appears on: /domains/cases/illinois-dcfs-augintel
EmpiricalCalifornia's tax agency ran the first state generative AI call-center assistant through a complete governed procurement …
California's tax agency ran the first state generative AI call-center assistant through a complete governed procurement arc: a competitive sandbox in which two vendors were each paid one dollar to test for six months in a secure environment, a 10-month pilot in a simulated environment, and production rollout to roughly 375 agents completed August 22, 2025 under a 12-month, 445,000 dollar contract. The pilot projected a minimum 1.5 percent per-call time saving, and by May 2026 the agency had declined to renew the contract, with the official who launched the project saying the working system 'didn't quite save as much time as we had hoped.'
Sources: stateofcaliforniagenaiportal2024, officeofthegovernorofcalifor2025, governmenttechnologyindustry2025, melhado2026
Appears on: /domains/cases/california-cdtfa-genai-call-center
EmpiricalThe benefit figures for the CDTFA call-center assistant are pilot projections from a simulated environment, self-reporte…
The benefit figures for the CDTFA call-center assistant are pilot projections from a simulated environment, self-reported by the agency through trade press and echoed in vendor marketing: the call-center chief said the simulated environment 'did show some potential to save at least one and a half percent time on some of those calls,' extrapolated to roughly 100,000 minutes per year and capacity for about 10,000 additional calls annually across roughly 800,000 yearly inquiries. No published production-environment measurement, methodology, or independent assessment report has been located.
Sources: governmenttechnologyindustry2025, symsoftsolutionsviabusinessw2025
Appears on: /domains/cases/california-cdtfa-genai-call-center
EmpiricalBy May 2026 CDTFA had declined to renew the 12-month, 445,000 dollar contract for its custom call-center assistant and s…
By May 2026 CDTFA had declined to renew the 12-month, 445,000 dollar contract for its custom call-center assistant and shifted to a similar off-the-shelf tool from Amazon Web Services under an existing larger contract, with Government Operations Secretary Nick Maduros saying the solutions 'worked in practice' but 'didn't quite save as much time as we had hoped.' The decision was a procurement authority action inside a broader portfolio review of eight executive-order-driven generative AI pilots totaling more than 5.8 million dollars, in which three other projects also ended while three continued.
Sources: melhado2026
Appears on: /domains/cases/california-cdtfa-genai-call-center
EmpiricalNew Jersey built and hosts its own generative drafting assistant for state employees, with a state-owned interface, host…
New Jersey built and hosts its own generative drafting assistant for state employees, with a state-owned interface, hosting and logs and a hosted commercial frontier-model service as the one external dependency in the serving path, an ownership arrangement that leaves the underlying model replaceable without reprocuring an application or moving employees onto a different product. The state reports that roughly 20,000 employees had used the tool across more than 300,000 sessions and more than 1,000,000 prompts by February 2026 at about one dollar per user per month, that access onboarding routes through a responsible-AI course whose curriculum is used by 25 or more states, and that unemployment-insurance staff rewriting claimant emails in plain language saw claimants respond 35 percent faster. Every outcome figure is state self-reported; the 35 percent figure has no published methodology and predates the statewide launch, and no inspector-general evaluation, external audit or peer-reviewed causal study of the tool was located.
Sources: njofficeofinnovation2026a, routefifty2025, statescoop2024, njofficeofinnovation2024, sofi2026
Appears on: /domains/cases/nj-ai-assistant
EmpiricalNew Jersey's joint policy circular 25-OIT-001 requires human review of all AI-generated content for accuracy, bias, comp…
New Jersey's joint policy circular 25-OIT-001 requires human review of all AI-generated content for accuracy, bias, completeness, accessibility and style; permits sensitive personal information only inside state-approved tools, naming the NJ AI Assistant, with Agency CIO approval; and requires State Chief Technology Officer clearance plus registration for resident-facing or decisional generative systems, a gate the staff-facing assistant has not been reported to pass. The same circular's text says all state employees 'should' take the responsible-AI course, a should-language obligation stated in the governing document itself, while the focal unemployment-insurance agency is reported to have reached full training and tool coverage of its staff. A statewide public-workforce survey preceded the November 2024 AI Task Force recommendations to the Governor, and user surveys and interviews drove the March 2026 rebuild, which added a visible reasoning section intended to help employees catch errors and hallucinations.
Sources: njofficeofinformationtechnol2025, stateofnewjerseyaitaskforce2024, njofficeofinnovation2026, routefifty2025
Appears on: /domains/cases/nj-ai-assistant
EmpiricalNew Jersey's unemployment-insurance and TDI/FLI teams built plain-language glossaries, reusable prompt libraries and qua…
New Jersey's unemployment-insurance and TDI/FLI teams built plain-language glossaries, reusable prompt libraries and quality-evaluation rubrics for AI-assisted translation into Spanish and Haitian Creole, with human review by professional translators, subject-matter experts and seven community organizations; the state reports a higher share of Spanish-language unemployment-insurance applications and better follow-through, without quantifying either. US Digital Response gave the department a 2025 SEED Award and republished the materials for reuse by other states, the associated responsible-AI curriculum is used by 25 or more states and local partners, and the March 2026 rebuild moved training guidance inside the tool itself.
Sources: newjerseydepartmentoflaboran2025, usdigitalresponse2025, routefifty2025, njofficeofinnovation2026, innovateus2024
Appears on: /domains/cases/nj-ai-assistant
EmpiricalThe Social Security Administration deployed a conversational question-and-answer chatbot on its national 800-number in A…
The Social Security Administration deployed a conversational question-and-answer chatbot on its national 800-number in April 2025, answering 74 frequently asked questions before a caller reaches an employee. Automation on that line went from roughly 300,000 handled calls a month in fiscal 2024 to roughly 2.9 million a month in fiscal 2025, peaking at 5.1 million automated calls in March 2025, while the agency served 68 million callers, a 65 percent increase over the prior year, with a workforce that fell about 10 percent net from roughly 57,000 to roughly 51,400 and about 1,000 field office employees reassigned onto 800-number duty. About 25 million fiscal 2025 calls ended with no service, abandoned in queue or met with a busy signal, and the busy rate spiked to 29.4 percent in March 2025 during the Social Security Fairness Act surge, which affected 3.2 million beneficiaries.
Sources: ssaofficeoftheinspectorgener2025, kffhealthnewsdariustahir2025, aarp2025
Appears on: /domains/cases/ssa-800-number-ai-assistant
EmpiricalThe agency's inspector general found the published telephone metrics arithmetically accurate as computed and documented …
The agency's inspector general found the published telephone metrics arithmetically accurate as computed and documented what they cover: the headline average speed of answer, 12.7 minutes in October 2024, a 29.7 minute peak in January 2025 and 7.0 minutes in September 2025, counts a caller who accepts a callback as a zero wait, and the roughly 25 million calls a year ending in a hang-up or a busy signal are excluded. The 23.8 million callers who accepted callbacks in fiscal 2025 waited an average of 61.5 to 151.8 minutes depending on the month, and the 9.3 million who declined and held waited an average of 18.8 to 100.1 minutes. A separate 87 percent first-contact-resolution figure comes from a post-call survey that is not offered to callers served only by automation, and no audit has published a misrouting or wrongful-disconnection rate for the chatbot itself.
Sources: ssaofficeoftheinspectorgener2025, nextgovfcw2025b
Appears on: /domains/cases/ssa-800-number-ai-assistant
EmpiricalDeployment decisions on this line have followed changes of leadership rather than published evaluation results. The chat…
Deployment decisions on this line have followed changes of leadership rather than published evaluation results. The chatbot was developed and tested under one administration and shelved as not ready, with the former chief information officer saying the team wanted to ensure the automation produced consistent and accurate answers and that this would take more time; the next administration deployed it in April 2025 and targeted extension to roughly 1,200 field offices by August 2025. A separate phone anti-fraud AI check was added in April 2025 and its three-day claim holds removed in mid-May 2025 after it flagged 2 of more than 110,000 claims while slowing retirement claim processing by 25 percent, a different tool whose figures do not merge with the chatbot's. The predecessor Next Generation Telephony Project, built by Verizon Business Network Services under a February 2020 contract, was abandoned on August 22, 2024 after about ten months and more than 160 million dollars paid, with an audit finding that contract lacked performance-based quality standards; the vendor of the current cloud platform is not named in the public audit. Five commissioners or acting commissioners served during fiscal 2025, and the metrics published on the agency's performance website were added and removed according to what each leadership believed were the most important metrics for the public.
Sources: kffhealthnewsdariustahir2025, nextgovfcw2025a, officeofsenatorelizabethwarr2025, ssaofficeoftheinspectorgener2025b, ssaofficeoftheinspectorgener2025
Appears on: /domains/cases/ssa-800-number-ai-assistant
EmpiricalCalifornia's Employment Development Department runs a two-tier chat assistant delivered under the Integrated Contact Cen…
California's Employment Development Department runs a two-tier chat assistant delivered under the Integrated Contact Center work stream of EDDNext, the state's roughly $1.258 billion modernization of its unemployment, disability and paid family leave systems. The unauthenticated public-site tier became available around the clock in the state's top eight working-age languages in May 2025 and served 554,792 unique customers across 2,103,782 messages between January 1 and June 30, 2025. A live agent chat channel for unemployment customers, which the department dates to July 2025, lets a customer escalate to a person on weekdays between 9 a.m. and 2 p.m. after identity verification, with account details passed to the agent, real-time machine translation in six non-English languages, a redacted transcript saved to the account and a post-chat survey. On May 8, 2026 a second, authenticated tier launched inside the customer portal, answering a signed-in unemployment customer's own claim status, payment and eligibility questions for claims filed in the past three years; the department reported more than 25,000 uses and nearly 18,000 fully self-service interactions in its first two weeks. The platform is documented as intent-based conversational AI; the public sources do not establish generative language modeling. All usage figures are agency self-reported, and the two tiers' counts belong to different systems.
Sources: californiaemploymentdevelopm2023, californiaemploymentdevelopm2025a, californiaemploymentdevelopm2025d, californiaemploymentdevelopm2025c, californiaemploymentdevelopm2026
Appears on: /domains/cases/california-edd-virtual-assistant
EmpiricalNo body oversees the EDD chat assistant as such: its accountability is inherited two levels up, from legislative and exe…
No body oversees the EDD chat assistant as such: its accountability is inherited two levels up, from legislative and executive scrutiny of the EDDNext programme it is a deliverable of. That scrutiny is substantial and has demonstrably changed programme behaviour — the 2026-27 Governor's Budget reverts $70.6 million of unused modernization funding early as a budget solution, and the core claims-system replacement was resequenced to do disability and paid family leave first with unemployment integration designated 'mandatory optional' to reduce risk for the state — and it is also documented as partial, since the adopted 2025-26 Budget Act retained extended spending authority against the Legislative Analyst's Office recommendation to drop it. What the oversight record engages with is budget, schedule and procurement: the 2026 analyst-office questions on the record concern the core project's vendor and the rationale for removing unemployment from it, and the February 2026 handout flags that new front-end functionality is linked to the legacy claims mainframe through 'informal and untested data bridges and custom-built interfaces' that have not been stress tested — a finding stated generically about new functionality, without naming the chat assistant. No inspector-general review, state-auditor evaluation or academic study of this assistant's answer accuracy or translation fidelity has been located in the public record.
Sources: californialegislativeanalyst2025, californialegislativeanalyst2025c, californialegislativeanalyst2026, californialegislativeanalyst2026a, californiastatesenate2026
Appears on: /domains/cases/california-edd-virtual-assistant
EmpiricalThe success measures published for the EDD chat assistant are deflection-shaped, and the department both produces them a…
The success measures published for the EDD chat assistant are deflection-shaped, and the department both produces them and reports them upward. Since July 2025 the agency states that 23 percent more customers resolve questions through self-service and 47 percent fewer need to speak with an agent after using it, alongside more than 2.1 million self-service actions since November 2024, more than 830,000 customers using callback since May 2024, more than 29,900 customers served by live-chat agents, and an in-chat issue-resolution rate above 80 percent. Those figures measure channel exit rather than verified problem resolution, and they are agency self-reported; the human channel they are measured against runs weekdays from 9 a.m. to 2 p.m. and reached nearly 6,000 unemployment customers a month as of October 2025. Separately, the integrator published on its own marketing blog that the bot stack reduced call volume by 35,000 calls a day, cut live-agent volume by 3,800 calls a day, saved constituents 684 hours daily and generated $8.4 million in annual savings; those are vendor claims with no independent verification.
Sources: californiaemploymentdevelopm2025b, californiaemploymentdevelopm2025c, intervisionsystems2025
Appears on: /domains/cases/california-edd-virtual-assistant
EmpiricalMassachusetts built a retrieval-grounded generative-AI Virtual Assistant in house in 83 days and launched it on motor-ve…
Massachusetts built a retrieval-grounded generative-AI Virtual Assistant in house in 83 days and launched it on motor-vehicle registry pages in April 2025 ahead of the May 7, 2025 federal identification deadline, describing it as a state-owned platform that replaced a vendor-managed rule-based chatbot; by 2026 the assistant's own support page listed motor-vehicle, toll, unemployment-assistance, tax, child-support, transitional-assistance, family-and-medical-leave and state-login content. The state reports more than 1,500 conversations a day and roughly 200,000 conversations since launch, a chat open rate rising from 1.29 to 2.92 percent, positive feedback rising from 10 to 56 percent, negative feedback down 60 percent, page abandonment down 40 percent, around-the-clock availability in English, Spanish and Portuguese, a knowledge base refreshed nightly from published state content, and a custom evaluation tool that flags answer issues within hours. Every one of those figures is an agency self-report; the technology agency withheld the cost and usage reports and the technical logs that would confirm its claims, and the underlying foundation model and hosting stack have not been publicly identified.
Sources: massachusettsdigitalservice2025, executiveofficeoftechnologys2026, executiveofficeoftechnologys2025, theshoestring2026, massachusettsdigitalservice2026
Appears on: /domains/cases/massgov-virtual-assistant
EmpiricalMassachusetts policy AI.001, effective January 31, 2025, requires human fact-checking of generative output, conspicuous …
Massachusetts policy AI.001, effective January 31, 2025, requires human fact-checking of generative output, conspicuous labeling of AI content, Chief Technology Officer approval for generative procurement and a generative-AI inventory, while routing privacy review through consultation with legal and security teams; the phrase privacy impact appears in it zero times. The Enterprise Privacy Office told the legislature in February 2025 that it has been using a Privacy Impact Assessment and was overseeing a pilot program with its risk and security teams to assess privacy risks during the contracting and development stages. Records obtained by an independent investigation showed at least forty AI use cases in the agency's internal survey with thirty-one withheld, and of the nine entries released not one recorded a completed privacy impact assessment, the fields left blank without explanation including for tools processing Social Security numbers and Medicaid data, with the agency spokesperson declining to answer questions about them; a count across all forty is therefore an inference the agency has not contradicted rather than a verified number, and the assessment is an internal office practice rather than a statutory mandate. After months of negotiation the Supervisor of Records ordered the agency to produce the withheld records for in camera review, allowing ten business days to produce and up to fifteen business days to review; no final determination was found as of July 20, 2026.
Sources: theshoestring2026, commonwealthofmassachusettse2025, executiveofficeoftechnologys2025, massachusettsexecutiveoffice2026
Appears on: /domains/cases/massgov-virtual-assistant
EmpiricalThe Massachusetts assistant answers from a knowledge base rebuilt nightly out of published state content, and the state …
The Massachusetts assistant answers from a knowledge base rebuilt nightly out of published state content, and the state reports that those published pages were updated and simplified to suit the assistant while chat analytics drive prompt tuning, content updates and enhancements, so the authoritative public statement of program rules is adapted to the channel that reads it. On the human-channel side, an independent institute brief reporting state figures records motor-vehicle registry calls falling by about 1,000 a day and emails by about 200 a week after launch, counts about twenty AI use cases publicly reported with three facing the public against forty in the internal survey, and recommends a formal public AI inventory and structured user feedback loops; the assistant offers three languages where the predecessor rule-based chatbot offered twenty. The state technology secretary committed publicly that there will be a human reviewing output before public distribution and framed AI as relieving stretched agencies without adding headcount, while a public-employee union representing about 9,000 state employees argued that staffing rather than AI tooling is what would let the transitional-assistance agency serve more clients.
Sources: massachusettsdigitalservice2025, pioneerinstitute2026, commonwealthbeacon2026, theshoestring2026, executiveofficeoftechnologys2025
Appears on: /domains/cases/massgov-virtual-assistant
EmpiricalMyFriendBen is an advisory multi-state benefits screener: an anonymous survey of about six minutes returns a report of p…
MyFriendBen is an advisory multi-state benefits screener: an anonymous survey of about six minutes returns a report of programs a household is likely eligible for, with estimated dollar values, and submits nothing, routing people instead to separate government application channels that make every determination independently. All of its eligibility math is computed by PolicyEngine, a separate nonprofit whose open-source codebase encodes federal and state statute; MyFriendBen deliberately built no proprietary rules logic, and the same substrate sits under each of its state deployments, so an encoding error would misestimate benefits in every state simultaneously and a correction would propagate to every state simultaneously. That substrate's scale is indicated by a $300,000 NSF POSE Phase I award announced on August 18, 2025. The shared-substrate structure is documented in both organizations' technical materials; no incident of a cross-state encoding error appears in the public record. The accuracy figure attached to it, 'over 90%', is a builder, infrastructure-organization and operator claim with no published methodology and no independent audit.
Sources: policyengine2025a, githubmyfriendbenorg2026, policyengine2025b, codethedreammyfriendbenncben2025
Appears on: /domains/cases/myfriendben
EmpiricalThe gap between benefits identified and benefits received is measured only from outside the system. Independent journali…
The gap between benefits identified and benefits received is measured only from outside the system. Independent journalism reported that in all of 2023 in Colorado about 5,500 households were screened, about $30 million in benefits was identified and about $5 million was estimated to have actually been obtained, with the median user reporting an income just over $8,200 a year and a household size of two. The deployment's own impact totals are self-published or funder claims and are mutually inconsistent, ranging across 20,000+ Coloradans, 55,000+ households, 100,000+ households, 115,000+ households nationally and 65,000+ families, with dollar figures of $33 million, $52 million, $58 million and $1.2 billion identified, on denominators that shift between identified, applied for, unlocked and delivered. The one commissioned evaluation, by the Urban Institute, is described by a funder's post as currently under way; its interim figures from 301 users cover discovery, intent and satisfaction rather than calculation accuracy or verified enrollment, and its primary report was not located in open search. The Aspen Institute's Financial Security Program finds that no systematic evaluation exists of whether benefits screeners increase benefits access.
Sources: thecoloradosun2024, deltafund2026, prnewswiremyfriendbencpalrel2026, aspeninstitutefinancialsecur2024
Appears on: /domains/cases/myfriendben
EmpiricalThe deployments are operated by independent local anchors rather than by one organization: Code the Dream with NC 211 in…
The deployments are operated by independent local anchors rather than by one organization: Code the Dream with NC 211 in North Carolina, Benefit Illinois's Illinois Benefit Hub, the MASSCAP community-action network in Massachusetts and Child Poverty Action Lab in Texas, each configuring its own program catalog on a shared white-label backend that carries per-state feature flags and admits a program only if it is worth at least $300 a year. The front-line users are employed by those partner agencies and others: 2-1-1 Colorado piloted the tool with its own call staff in autumn 2023, and 100 to 130 individuals at one Dallas employer use it weekly, with one partner reporting that it serves about 4,000 individuals a year through it. The builder's stated reason for navigator uptake is that training a case manager on a single program's rules otherwise takes months. Screening data has also flowed outward into policy: a Colorado Child Tax Credit Calculator derived from the same code identified $177 million, and the builder credits the surrounding ecosystem with unlocking an $810 million state child tax credit through HB24-1311, a rule the shared substrate then had to encode. The catalog, threshold, uptake and policy-loop figures are builder, operator and joint-release claims.
Sources: garycommunityventures2025, prnewswiremyfriendbencpalrel2026, githubmyfriendbenorg2026, masscapmassachusettsassociat2026
Appears on: /domains/cases/myfriendben
EmpiricalSafeRent Solutions (formerly CoreLogic Rental Property Solutions) sold landlords a Registry ScorePLUS tenant-screening m…
SafeRent Solutions (formerly CoreLogic Rental Property Solutions) sold landlords a Registry ScorePLUS tenant-screening model returning a single 200-800 'lease performance risk' score plus an accept/decline/conditional recommendation measured against a landlord-chosen cutoff (500 in the complaint's worked example, with an optional 450-499 conditional band), built from credit bureau reports and scores including non-tenancy debt, bankruptcy records, past-due accounts, payment performance, and eviction and landlord-tenant court records, weighted 'according to their statistical significance in predicting lease performance'; factor weights were undisclosed to landlords, applicants and the public, and the company's marketing told buyers a landlord cannot change the screening algorithm. The plaintiffs' pleaded theory concerns a missing input: as of 2021, HUD data cited in the complaint shows Black and Hispanic voucher holders in Massachusetts paying an average of $423/month toward rent and utilities while public housing authorities paid landlords an average of $1,159/month directly (at least 73.26% of the expected payment), on tenancies averaging over 21 years in the same unit, and that subsidy is not a model input. These are plaintiff-pleaded figures accepted by the court as allegations at the motion-to-dismiss stage only.
Sources: louisv2022, louisv2023, findlawcaselaw2023
Appears on: /domains/cases/saferent-score-voucher-screening
EmpiricalThe authority over a SafeRent-screened tenancy decision was documented as inverted: the vendor authored the weights and …
The authority over a SafeRent-screened tenancy decision was documented as inverted: the vendor authored the weights and returned the answer without any housing relationship to the applicant, property management picked the cutoff 'in consultation with SafeRent' without knowing how scores are computed, and the leasing desk that signed the lease received only the bottom-line score ('The Leasing Manager does not receive the detailed credit information at the time of running the applicant screening'; 'CoreLogic sends us a number, and if it is above the predetermined approved number, we move forward... We do not know why they were denied') and disclaimed override authority to a named plaintiff in writing ('we do not accept appeals and cannot override the outcome of the Tenant Screening'). On July 26, 2023 Judge Angel Kelley denied the motions to dismiss the FHA 3604(a),(b) and Massachusetts c.151B race and source-of-income claims on the reasoning that SafeRent 'effectively controls' approval decisions because it alone builds and conceals the algorithm, while dismissing the c.93A consumer counts. The single documented pre-settlement reversal ran outside the screening system entirely: a July 22, 2021 denial, a personal appeal rejected August 26, and a second appeal drafted by the tenant-advocacy organization City Life/Vida Urbana delivered September 1-8, 2021 — roughly six weeks of organized advocacy for one reversal, from which no reversal rate can be generalized.
Sources: louisv2022, louisv2023, unitedstatesdepartmentofjust2023a, civilrightslitigationclearin2025
Appears on: /domains/cases/saferent-score-voucher-screening
EmpiricalThe Louis v. SafeRent settlement (executed March 28, 2024; finally approved November 20, 2024; $2.275M total, $1.175M fu…
The Louis v. SafeRent settlement (executed March 28, 2024; finally approved November 20, 2024; $2.275M total, $1.175M fund with no reversion) remedied the case on the output channel rather than in the model: for five years SafeRent may return no SafeRent Score, no other tenant screening score and no accept/decline recommendation on a report for a voucher-holder application, providing a report of underlying information instead (3.5.2); for the 'market' and 'no-credit' products the score is suppressed by default unless the landlord affirmatively certifies the applicant is not a voucher recipient (3.5.3); any tenant screening score may re-enter the voucher channel only if 'found to be valid when used for voucher-holders by the National Fair Housing Alliance' or another organization mutually agreed by class counsel and SafeRent (3.5.5(ii)(1)); third-party credit scores (FICO, VantageScore) may still be passed through with source disclosure (3.5.5(iii)); customer training is required (3.5.4); and the court retains continuing exclusive enforcement jurisdiction for five years from SafeRent's certification (6.9), approximately through 2029-2030. The geographic reach of the practice-change terms is ambiguous: 3.5.2 and 3.5.3 are not expressly limited to Massachusetts, while the classes, the notice population and the plaintiffs' framing are Massachusetts-centered, and some coverage characterizes the changes as company-wide. Suppression therefore shifts rather than eliminates the decision input. SafeRent settled without admitting liability and states it 'continues to believe the SRS Scores comply with all applicable laws' (vendor claim); as of 2026-07-20 no public post-settlement compliance report, enforcement motion or validating-organization event had been located, so the gate exists on court-approved terms and has never been observed operating. The >18,000 figure is Massachusetts applications scored below the housing provider's accept threshold between May 25, 2020 and September 27, 2023 (the records-pull window), an upper-bound flow figure that does not identify voucher status or race, paired with SafeRent's own October 2023 estimate of 3,300-4,200 putative class members.
Sources: louisv2024, epiqclassactionservices2024, associatedpressviafortune2024, greaterbostonlegalservices2024
Appears on: /domains/cases/saferent-score-voucher-screening
EmpiricalThe United States and ten plaintiff states allege in United States v. RealPage, Inc. (M.D.N.C., filed 23 August 2024, am…
The United States and ten plaintiff states allege in United States v. RealPage, Inc. (M.D.N.C., filed 23 August 2024, amended 7 January 2025) that competing landlords contractually fed a single vendor nonpublic, competitively sensitive information — executed new-lease rents, renewal offers and rates, lease terms and occupancy signals — to train and run a common pricing algorithm that recommended rents back to all of them: the complaint alleges at least 80 percent of the commercial revenue-management software market for multifamily housing and data agreements reaching over 16 million units nationwide including units of landlords who were not customers, describes an 'Auto Accept' setting that implemented daily recommendations with no human review, a 'Governor' feature alleged to constrain price decreases more than increases, and vendor pricing advisors who monitored client acceptance and pushed property managers toward compliance, and reports a national acceptance rate of 40 to 50 percent for new leases across January 2017 to June 2023 against internal analysis finding nearly 60 percent of final floor-plan prices within 2.5 percent of the recommendation and more than 85 percent within 5 percent — allegations only, with no defendant having admitted wrongdoing and no liability adjudicated.
Sources: usdepartmentofjusticeantitru2024, usdepartmentofjusticeofficeo2024, paul2025
Appears on: /domains/cases/realpage-rent-algorithm
EmpiricalThe consent decree the Department of Justice proposed on 24 November 2025 edits the deployment's structure rather than i…
The consent decree the Department of Justice proposed on 24 November 2025 edits the deployment's structure rather than its accuracy: competitor data used in models must be at least 12 months old and not drawn from active leases, the existing demand and supply models may not be trained with a geographic variable narrower than a state, the 'Governor' feature must treat increases and decreases symmetrically, automatic-acceptance ranges must be user-set and off by default, vendor-hosted meetings of competing landlords are barred, and a court-appointed monitor holds sweeping oversight for three years under a seven-year term with inspection rights, a written antitrust compliance program and a cooperation obligation — with no fine and no admission of liability; the Proposed Final Judgment and Competitive Impact Statement were published on 5 December 2025 (90 FR 56286), the stipulation was entered 26 March 2026, and the Department responded to eight public comments on 8 May 2026, with final public-interest entry by the court still pending as of July 2026 and no compliance findings published by the monitor.
Sources: usdepartmentofjusticeofficeo2025b, federalregister2026, hoganlovells2025, wilsonsonsini2025
Appears on: /domains/cases/realpage-rent-algorithm
EmpiricalThe White House Council of Economic Advisers estimated in December 2024 that rental pricing algorithms cost United State…
The White House Council of Economic Advisers estimated in December 2024 that rental pricing algorithms cost United States renters more than $3.8 billion in 2023, roughly $70 per month per unit in algorithm-priced buildings, with the software pricing at least 10 percent of all US rental units and nearly one in four multifamily rental units — model-based counterfactual estimates the council explicitly framed as a lower bound because non-participating landlords also raised rents in response, and never measured overcharge; separately, the Middle District of Tennessee preliminarily approved 26 private settlements involving 27 landlord defendants totalling $141.8 million on 21 November 2025 for a class of renters of covered properties between 18 October 2018 and 21 November 2025, over objections from five state attorneys general, with a second batch of roughly $218 million announced in 2026.
Sources: whitehousecouncilofeconomica2024, americanbarassociationantitr2025, multifamilydive2026
Appears on: /domains/cases/realpage-rent-algorithm
EmpiricalNew York City's Homebase homelessness-prevention program has since June 2012 routed applicant households on a 15-item Ri…
New York City's Homebase homelessness-prevention program has since June 2012 routed applicant households on a 15-item Risk Assessment Questionnaire scored 0-25 (each answer worth 1-3 points), distilled by backward elimination from a Cox proportional-hazards model of shelter entry fitted to 11,105 families who applied October 2004 - June 2008 (12.8% entered shelter within three years; decile risk 1% to 37%); a total at or above 7 points routed a household to 'full' services (financial assistance, case management, legal and mediation referrals) rather than a 'brief' one-or-two-visit contact. It is a regression-derived additive point screener, not a machine-learning system: the city's statutory algorithmic-tool register files it under computation type 'Scoring', purpose 'Resource allocation', autonomy 'Monitored', frequency 'Daily', with 'Vendor(s): None' in the CY2025 entry. Against caseworker judgment, which had deemed 66.5% of applicants eligible, the instrument would have increased correct targeting of families entering shelter by 26% and cut misses by almost two-thirds at equivalent false-alarm rates. Services are delivered by seven contracted nonprofit providers across 26 neighborhood offices, and The Department of Homeless Services states the network serves more than 25,000 at-risk households a year. The causal effect evidence for the program comes from outside the deploying agency's own research office: a randomized controlled trial by Abt Associates, commissioned by the Department of Homeless Services (2010-2013, 295 families with children analyzed across eleven sites), found the share spending at least one night in shelter falling from 14.5% to 8.0%, the share applying for shelter falling from 18.2% to 9.3%, and average shelter nights falling by 22.6; an independent community-district difference-in-differences study estimated roughly 5-11% fewer family shelter entries (11.2 log points, 95% CI 3.6-18.8), a $14.2M annual budget avoiding an estimated $20-44M of shelter expenditure. The separately claimed 'prevention rate' of around 90-97% is a city performance metric with no counterfactual and is not an effect size. The threshold of 7 describes the documented pre-2023 configuration; no public source states the threshold in force after the 2023 item revision.
Sources: shinnm2013, nycofficeoftechnologyandinno2026, rolstonh2013, goodmans2016, instituteforchildrenpovertya2024, newyorkcitydepartmentofhomel2025
Appears on: /domains/cases/nyc-shelter-entry-prediction
EmpiricalThe deploying agency's own Office of Research & Policy Innovation published a peer-reviewed re-examination (Housing Poli…
The deploying agency's own Office of Research & Policy Innovation published a peer-reviewed re-examination (Housing Policy Debate, 2022) of 48,450 deduplicated families with children applying 2013-2016 (58,674 family-years, over a period in which the caseload rose from about 600 cases a month in 2013 to more than 1,500 a month in some 2015 months), and documented two distortions against itself. First, score clustering at the eligibility cutoff: 4,269 cases scored 6, 10,634 scored exactly 7, and 7,756 scored 8, which its researchers wrote 'suggests that Homebase staff may be focused on getting families to that threshold so they qualify for full services' — while the agency separately asserts that the override valve reduces incentives for workers to misreport data to ensure eligibility. Second, override outcomes: 6.1% of the 58,674 applications departed from the score with mandatory supervisor approval (4.8% up to full services, 1.4% down to brief), and only 3.7% of the families moved up later applied to shelter against 25.8% of the families moved down, the paper concluding that worker judgment is less accurate than the RAQ on average; a ten-case note review attributed many downward decisions to needs beyond the program's capacity, such as families needing an apartment immediately with no funds, rather than to a judgment about risk. The same paper reported 73.9% of applications at or above the threshold, shelter application within two years at 13.7% above the cutoff against 5.9% below it (chi-square 699.98, p<.001), an area under the curve of 0.7387 for the deployed instrument against 0.7389 for the revised one, and a simulated alternative threshold of 5 raising precision from 13.7% to 15.2% at similar enrollment volume. These are agency-authored figures, published in a peer-reviewed venue; the paper's own caveats are that the observed score distribution is confounded by the clustering it documents and by low-scoring self-selection out of intake, and that outcomes compared by service tier are contaminated by the overrides being measured.
Sources: mullenej2022, newyorkcityofficeoftechnolog2024, farrelldc2023
Appears on: /domains/cases/nyc-shelter-entry-prediction
EmpiricalLocal Law 35 of 2022 (passed by the NYC Council 2021-12-15, lapsed into law unsigned 2022-01-14) added Admin. Code sec. …
Local Law 35 of 2022 (passed by the NYC Council 2021-12-15, lapsed into law unsigned 2022-01-14) added Admin. Code sec. 3-119.5, requiring every city agency to report by 31 December every algorithmic tool it used one or more times during the prior calendar year — expressly including tools that 'generate risk scores' or 'determine what resources are allocated to particular groups or individuals' — with six mandatory disclosure elements, compiled by the Office of Technology and Innovation into a public report delivered to the mayor and Council speaker each 31 March. DSS filed the Homebase RAQ in all four cycles to date (CY2022-CY2025), each entry disclosing in the agency's own words the June 2012 start date, the 2004-2008 training data analyzed with academic researchers, the input factors, the points-and-threshold eligibility mechanism and the worker-override-with-supervisor-permission rule. The loop demonstrably fired: after initial publication of the CY2023 report the register carries the note 'Update 3/27/2024 - The Department of Social Services updated their reporting to include changes made to the Homebase Risk Assessment Questionnaire after initial publication of the report', disclosing the 2023 item revision and adding its peer-reviewed citation; the Council's Committee on Technology held an oversight hearing on the regime on 2024-10-28, at which OTI described coordinating 45 agencies plus 24 further offices. Register entries are unaudited agency self-reports. Separately, three government audits have examined this program and none examined the risk model: the city comptroller's MG12-125A (2013-06-27) found no written monitoring policies, no records of initial ineligibility determinations and all provider risk assessments announced in advance, with the unannounced-visit recommendations rejected; the city comptroller's January 2020 audit of HRA's oversight of the then-$53M-a-year program found 80 of 240 required provider case-file reviews performed, 2,661 of 24,938 FY2018 households (11%) returning one to four times within twelve months, $2,271,797 in provider advances unrecouped some sixteen months after closeout, and 5 of 28 visited client homes not habitable (4 never fixed), issuing 19 recommendations; and the state comptroller's audit 2023-N-8 (issued 2026-01-07, scope July 2021 - July 2025) examined the downstream CityFHEPS rental-subsidy channel for DSS Homebase clients (57,888 new cases and 123,762 individuals housed since 2018 through March 2025; spending $176M in FY2019 rising to $834M in FY2024) and found units with hazardous violations approved, 30 of 75 sampled case records without evidence of income verification, and rents averaging $525 a month above comparables in eleven of thirty sampled statewide cases. None of the three may be cited as an audit of the algorithm.
Sources: newyorkcitycouncil2022, newyorkcityofficeoftechnolog2024, newyorkcitycouncilcommitteeo2024, officeofthenewyorkcitycomptr2013, officeofthenewyorkcitycomptr2020, officeofthenewyorkstatecompt2026
Appears on: /domains/cases/nyc-shelter-entry-prediction
EmpiricalCrimSAFE, a criminal-record tenant-screening product sold by CoreLogic Rental Property Solutions, computes no risk score…
CrimSAFE, a criminal-record tenant-screening product sold by CoreLogic Rental Property Solutions, computes no risk score: it is a deterministic record-matching and filtering engine that matches applicant identity data against a database of court and arrest records aggregated from more than 800 US jurisdictions, classified into three primary categories (Crimes Against Property, Crimes Against Persons, Crimes Against Society) with sub-classifications, and applies filter criteria the HOUSING PROVIDER configures: offense type, disposition and severity across felony and non-felony convictions and charges, and a lookback period configurable from 0 to 99 years for convictions and 0 to 7 years for charges (federal consumer-reporting law permits reporting non-conviction records for seven years). The output is a report carrying a lease decision driven by the provider's criteria plus a credit score, a Record(s) Found flag, message text the provider authors, full record detail for the users the provider authorizes, and an optional provider-customizable adverse-action letter template. Every new CrimSAFE user is by default authorized to receive full record data, with no cap on how many users get full access; a provider must affirmatively change configuration settings to restrict full reports to senior managers. In April 2016 Carmen Arroyo's application to move within ArtSpace in Windham, Connecticut, so her son Mikhail could live with her after a 2015 injury, was denied on 26 April 2016 after the screen returned Record(s) Found; the only matched record was a pending Pennsylvania shoplifting charge, later withdrawn in April 2017. Mikhail's report carried the vendor's default message: 'Please verify the applicability of these records to your applicant and proceed with your community's screening policies.'
Sources: connecticutfairhousingcenter2026, connecticutfairhousingcenter2023, nationalhousinglawproject2018
Appears on: /domains/cases/corelogic-crimsafe
EmpiricalThe location of decision authority in this deployment was the contested question, and the Second Circuit resolved it com…
The location of decision authority in this deployment was the contested question, and the Second Circuit resolved it component by component on 20 February 2026 (Cabranes, Wesley, Menashi; opinion by Menashi; Nos. 23-1118(L), 23-1166(XAP)). WinnResidential had suppressed full reports from its own on-site staff so that leasing decisions involving criminal records would be made 'by someone in a more elevated position,' out of concern about leasing-commission incentives, so the on-site agent saw only a Record(s) Found flag and told Arroyo the application was denied without individualized review; answering the state commission's complaint, WinnResidential said it did not know 'the facts behind the criminal background findings' because it had 'trust' in CoreLogic's reports. After a ten-day bench trial Judge Vanessa L. Bryant ruled on 20 July 2023 that CrimSAFE does not disqualify applicants because the housing provider decides what records matter and whether to deny; the earlier August 2020 summary-judgment characterizations that the companies 'acted hand-in-glove' and that CoreLogic 'was an integral participant' belong to that posture and did not survive trial. The Second Circuit rejected the threshold reasoning that screening companies sit categorically outside the Fair Housing Act — a point the United States had urged as amicus on 24 November 2023 — but affirmed on proximate cause, holding the denial came after a chain of the provider's discretionary decisions (configuration, record relevance, staff access, adverse-action letters, final approval) and quoting the district court's own sentence with approval: 'No housing provider who uses CrimSAFE could reasonably believe that CoreLogic makes housing decisions for them.' It rejected the 'cat's paw' theory because the screening policies applied were the provider's own, rejected liability for failing to restrict lawfully reportable non-conviction records as extending liability beyond the first step, and dismissed the Connecticut Fair Housing Center's own claim for lack of Article III standing under a 2024 organizational-injury doctrine unrelated to tenant screening, vacating rather than deciding its merits. Disparate impact was never proven: the race and national-origin claim failed at the prima facie causation step, so the underlying statistics were never adjudicated.
Sources: connecticutfairhousingcenter2026, connecticutfairhousingcenter2023, connecticutfairhousingcenter2020, unitedstatesdepartmentofjust2023b
Appears on: /domains/cases/corelogic-crimsafe
EmpiricalThe correction loop in this deployment ran through records the subject could not see at the point of harm. CoreLogic's A…
The correction loop in this deployment ran through records the subject could not see at the point of harm. CoreLogic's Authentication Procedure Guide listed only a notarized power of attorney as third-party authorization and escalated 'any scenarios not covered' to a supervisor; staff demanded a power of attorney that Mikhail Arroyo, a conservatee, was legally incapable of executing — a demand the trial court called an 'impossible condition' — across a blocked window running from the 24 June 2016 request with a conservatorship certificate to mid-November 2016, when a 1 November call escalated to CoreLogic's legal department and the company agreed about two weeks later that a conservatorship certificate with a visible probate seal would suffice; the resubmitted copy again lacked a visible seal and the disclosure was never completed. That escalation is the only documented change in vendor behavior in the record. The family learned which record had caused the denial in December 2016, roughly eight months after the denial, from the housing provider rather than the vendor, and the adverse-action letter that should have triggered the correction loop was sent but never received; correction ultimately happened at the original source when Arroyo petitioned the Pennsylvania court and the charge was withdrawn in April 2017. Arroyo's was the first and only conservator file-disclosure request CoreLogic had ever received (478 F. Supp. 3d at 282), a finding both courts used to defeat the disability disparate-impact claim. The Connecticut Commission on Human Rights and Opportunities held an evidentiary hearing on 13 June 2017 and WinnResidential approved the move-in ten days later — about fourteen months after the denial, and the only oversight action documented to have changed an outcome. The sole liability finding, a willful consumer-reporting violation carrying $1,000 statutory and $3,000 punitive damages awarded in July 2023, was reversed on 20 February 2026; final vendor liability was zero, and no injunction, consent decree or policy change was ever ordered.
Sources: connecticutfairhousingcenter2026, connecticutfairhousingcenter2020, connecticutfairhousingcenter2023, courtlistenerandthefreelawpr2023
Appears on: /domains/cases/corelogic-crimsafe
EmpiricalNew York City launched NYC Teenspace on November 15, 2023 as a three-year, $26M contract buying population access to Tal…
New York City launched NYC Teenspace on November 15, 2023 as a three-year, $26M contract buying population access to Talkspace's existing consumer teletherapy platform for all city residents aged 13-17, with unlimited asynchronous messaging (therapist replies five days per week) plus one 30-minute live session per month from NY-licensed clinicians; inside the therapy chat a proprietary NLP model scans teen-authored messages for suicide and self-harm language and fires an urgent real-time alert to the teen's own treating therapist, which never acts autonomously — escalation (child-protective-services referral, intensive therapy, hospitalization) is entirely the clinician's; in the first six months the scan flagged roughly 50 teens as moderate-to-high suicide risk (under 1% of about 6,800 users) while therapists navigated 36 high-risk events including suicide attempts, a child-abuse report and a drug overdose; enrollment ran from about 6,800 (May 2024) to 19,000+ (DOHMH letter, Dec 1 2024) to a vendor-reported 45,000+ (Feb 2026), with about 80% of early registrants Black, Hispanic, AAPI, bi-racial or Native American, about 70% female, and more than half resident in the city's priority TRIE neighborhoods.
Sources: nycofficeofthemayor2023, chalkbeatnewyork2024, newyorkcitydepartmentofhealt2024, newyorkcitydepartmentofhealt2026
Appears on: /domains/cases/nyc-teenspace-talkspace
EmpiricalA September 2024 advocacy tracker audit of the NYC Teenspace pages counted 15 ad trackers and 34 cookies sharing teen vi…
A September 2024 advocacy tracker audit of the NYC Teenspace pages counted 15 ad trackers and 34 cookies sharing teen visitors' personally identifiable information with recipients including Facebook/Meta, Amazon, Google and Microsoft, alongside sensitive intake data (name, date of birth, address, school, gender, mental-health screening answers) collected before parental consent was secured; DOHMH's December 18, 2024 letter (General Counsel Landau, Chief Privacy Officer Elcock) recorded that Talkspace removed all social-media and advertising trackers as of December 11, 2024, that the sign-up flow was minimized to age and zip code only, and that the contract's Data Security Rider would be amended to ban marketing use and trackers outright; advocates then documented continuing tracker-based disclosure on pages beyond the landing page (including the page hosting the revised privacy policy) on January 9 and February 12, 2025 with no DOHMH response to their follow-up letter after more than a month, and February 2025 technical testing found the NYC landing page clean as of January 24 while Talkspace's Seattle and Baltimore teen pages still transmitted visitor IP addresses to TikTok, Meta, Snapchat, Google, X, Reddit, LinkedIn, Spotify and Quora until the reporter's inquiry — several recipients being companies NYC had sued in February 2024 over teen mental-health harm. Talkspace's Chief Privacy Officer stated no personal medical information was transmitted; advocates dispute the completeness of the remediation, so the leakage is identifier-level and contested-scope.
Sources: parentcoalitionforstudentpri2024, newyorkcitydepartmentofhealt2024, parentcoalitionforstudentpri2025, gizmodo2025
Appears on: /domains/cases/nyc-teenspace-talkspace
EmpiricalThe only published performance figure for the Teenspace suicide-alert algorithm is Talkspace's own claim of 83% accuracy…
The only published performance figure for the Teenspace suicide-alert algorithm is Talkspace's own claim of 83% accuracy versus a human expert, resting on a 2020 Psychotherapy Research study of ADULT platform data rather than any measurement of the teen population it reads, and the independent evaluation DOHMH said in a September 12, 2024 email that it was planning had still not been published as of mid-2026, with no OIG, comptroller or FTC action specific to Teenspace located; outcome figures (65% 'reported improvement' at six months; 66% 'measurable clinical improvement' among 45,000+ enrollees) are vendor or city self-reports with undisclosed instruments and no independent verification; meanwhile the record store the program writes into is described by Talkspace's CEO as 8 billion words, 140 million messages and 6.2 million assessments, is subpoenable (a Talkspace user's complete therapy history was obtained by her former employer and used in court in 2026 — an adult employer-benefit member, not a Teenspace teen), is stated by the vendor to be used for training behavioral-health LLMs, and transfers intact under the pending $835M Universal Health Services acquisition announced March 9, 2026 and expected to close in Q3 2026, while the NYC contract is still live.
Sources: talkspace2023, kdive2024, proofnews2026, prnewswireanduniversalhealth2026
Appears on: /domains/cases/nyc-teenspace-talkspace
EmpiricalCharacter.AI built its crisis-safety stack in four dated increments, each within days to weeks of a specific external ev…
Character.AI built its crisis-safety stack in four dated increments, each within days to weeks of a specific external event: a self-harm phrase screen referring users to the 988 Lifeline (vendor blog dated 22 October 2024, the day the Garcia wrongful-death complaint was filed in M.D. Fla., publicized the 23rd), a separate more restrictive under-18 model (December 2024, after two Texas family suits and a 15-company Texas Attorney General SCOPE Act investigation), a weekly parental usage summary carrying time spent and top characters while deliberately excluding chat content (25 March 2025), and — one week after an investigation found dozens of harmful personas live and seven weeks after FTC 6(b) orders reached seven companies — the removal of open-ended chat for under-18 users, announced 29 October 2025 and effective 25 November, with an interim two-hour daily cap ramping down and a two-stage age-assurance stack; on 7 January 2026 the company, both founders, and Google agreed to settle the Garcia case and four others in New York, Colorado, and Texas, terms undisclosed and approval still required, so causation was never adjudicated. Each retrofit is a company authority action temporally associated with an external event; the operator's own announcement cited several pressures at once, and no sole cause is asserted.
Sources: characterai2024, characterai2025a, cnnbusiness2026
Appears on: /domains/cases/character-ai-crisis-response
EmpiricalThe only direct measurement of Character.AI's self-harm detector is a single adversarial test published 29 October 2024:…
The only direct measurement of Character.AI's self-harm detector is a single adversarial test published 29 October 2024: across sixteen conversations with mental-distress-focused bots the 988-hotline pop-up appeared three times, triggered by two exact phrasings ('I am going to commit suicide' and 'I will kill myself right now') while many equally explicit statements including 'I want to end my life' did not trigger it, dismissable and non-blocking so the chat continued, with firing frequency increasing after the publication contacted the company — a sample of sixteen, and therefore anecdotal rather than a measured rate. No independent evaluation of the sensitivity or specificity of the crisis classifier, the two-stage age-assurance classifier, or the separate under-18 model exists in either direction; every efficacy claim about the retrofits is a vendor claim. The first compelled quantitative record of this detector class is prospective: California SB 243, signed 13 October 2025 and operative 1 January 2026, requires a crisis protocol issuing referral notifications and, from 1 July 2027, an annual report to the state Office of Suicide Prevention of the number of notifications issued, posted publicly — a count of referrals, not a measure of what they achieved.
Sources: futurism2024, californialegislativeinforma2025
Appears on: /domains/cases/character-ai-crisis-response
EmpiricalCharacter.AI's crisis pathway terminates in an automated referral to an external hotline with no staff review, no escala…
Character.AI's crisis pathway terminates in an automated referral to an external hotline with no staff review, no escalation call, and no capture of what followed; the human roles the record does document sit elsewhere — trust-and-safety moderators taking down user-authored characters reactively, parents receiving a weekly usage summary with no transcript access and only where the teen adds them, and, after November 2025, third-party identity reviewers adjudicating contested age determinations. The external review tier is by contrast unusually crowded and repeatedly behavior-forcing: two federal district courts, the Texas Attorney General twice, FTC 6(b) compulsory orders to seven companies on 11 September 2025, California SB 243, and a Senate subcommittee hearing, alongside investigative journalism that on 22 October 2025 — a year into the litigation — found dozens of harmful bots live including a 'Bestie Epstein' persona that had logged almost 3,000 chats, a gang simulator, school-shooter personas, and a 'doctor' giving antidepressant-tapering instructions. None of these reviewers ever operated inside the decision loop; what they moved was the structure itself.
Sources: thebureauofinvestigativejour2025, ftc2025b, characterai2025
Appears on: /domains/cases/character-ai-crisis-response
EmpiricalFour days after a labor board certified its paid helpline staff's union election, a national eating-disorder nonprofit t…
Four days after a labor board certified its paid helpline staff's union election, a national eating-disorder nonprofit told those staff they were terminated and that a chatbot would replace a helpline that had fielded nearly 70,000 contacts in 2022 with six paid staff, about two supervisors, and up to roughly 200 trained volunteers; the chatbot was suspended on May 30, 2023, two days before it was to become the sole channel, after testers published screenshots of it recommending a 500 to 1,000 calorie daily deficit, a 1 to 2 pound weekly loss, a 2,000 calorie cap, regular weigh-ins, and where to buy skinfold calipers, and the helpline closed on June 1 as scheduled while the chatbot was already offline.
Sources: wells2023, kffhealthnews2023, nprshots2023, picchi2023, npr2023a
Appears on: /domains/cases/neda-tessa-chatbot
EmpiricalThe deployed chatbot had two layers with different evidence status: a closed, pre-scripted program the vendor and the re…
The deployed chatbot had two layers with different evidence status: a closed, pre-scripted program the vendor and the research team describe as unable to depart from its authored content, and a generative question-and-answer feature the operating vendor added in what its chief executive called a systems upgrade covered by the client's contract, a reading the client's chief executive denies by saying the organization was never advised of the changes and would not have approved them; the vendor's public account of the harmful outputs was mixed, saying it was still trying to determine how a closed system allowed such content, so the attribution of the May 2023 advice to the generative layer is the reported and attributed explanation rather than an established mechanism, and the contract scope stays contested with no adjudicated breach.
Sources: nprshots2023, kffhealthnews2023, thewraprepublishedonyahoo2023
Appears on: /domains/cases/neda-tessa-chatbot
EmpiricalThe earliest documented external warning about harmful chatbot responses came in October 2022 from the executive directo…
The earliest documented external warning about harmful chatbot responses came in October 2022 from the executive director of a peer eating-disorder organization, and the specific language she flagged was quickly removed after she reported it, with no documented systemic review, output monitoring, or reassessment of the deployment following; the vendor's chief executive said the flagged language was part of the pre-scripted content rather than the generative layer, which the research team denies, so its provenance is contested, and system-level action arrived roughly seven months later when public screenshots circulated.
Sources: nprshots2023
Appears on: /domains/cases/neda-tessa-chatbot
EmpiricalThe Veterans Health Administration built an opioid overdose and suicide risk model in-house, with no commercial vendor, …
The Veterans Health Administration built an opioid overdose and suicide risk model in-house, with no commercial vendor, on its own electronic health record and Corporate Data Warehouse: fitted on 1,135,601 patients with an opioid prescription in fiscal 2010 against 23,790 overdose-related or suicide-related events among them in fiscal 2011 (a 2.1% base rate), reporting an area under the curve above 0.80 in training and test sets, and refreshed nightly as a continuous one-year risk estimate binned into percentile tiers on a population-management dashboard that shows each patient's risk factors and the guideline-recommended mitigation actions rather than a bare number. Its predictors are administrative - demographics, pharmacy records including opioid type and dose and co-prescribed sedatives, mental health and substance use disorder diagnoses, prior overdose-related and suicide-related events, detoxification episodes, and emergency department and other utilization history - so no structured risk questionnaire and no clinician-scored instrument feeds the score. VHA Notice 2018-08, issued 8 March 2018 and effective 18 April 2018, required all 140 VHA medical centers to convene interdisciplinary teams and case-review every patient in the very-high-risk tier, initially the top 1% of scores at a threshold of 0.166. The mandate compels an assessment of risk, prescription appropriateness and mitigation options and never an action: there is no mandated taper and no automatic prescription cutoff tied to the score, and every clinical decision remains with the treatment team. Across 44,042 patients in the top 1% to 5% band from April 2018 to March 2020 the mandate raised the odds of receiving a case review 5.1-fold (95% CI 3.64-7.23) and added 0.498 risk-mitigation strategies per patient (95% CI 0.39-0.61). Facility completion had a median of 71% (interquartile range 48-95%), about one facility in five met the 97% target, and each of the 89 surveyed facilities used a median of 23 distinct implementation strategies. Team composition and workflow were left to each facility.
Sources: oliva2017, strombotne2023, minegishi2022, rogal2020
Appears on: /domains/cases/va-storm-opioid-risk
EmpiricalThe Veterans Health Administration randomized two features of its own opioid case-review policy across its 140 medical c…
The Veterans Health Administration randomized two features of its own opioid case-review policy across its 140 medical centers, and the two randomizations are separate experiments whose findings do not combine. The first was timing: when a facility's mandated review tier widened from the top 1% of risk scores to the top 5%, executed as a stepped wedge with waves on 12 February 2019 and 13 August 2019, which the primary trial paper places at study months 11 and 17. The second was language: whether a facility's copy of the policy notice carried an accountability paragraph naming a 97% case-review completion target, with quarterly reporting to the national Office of Mental Health and Suicide Prevention and technical assistance and action plans for facilities below it. Seventy facilities received that paragraph and seventy did not. Across 16,272 very-high-risk patients (8,734 in the accountability arm, 7,538 outside it), about 57% received a case review overall against a pre-mandate baseline of 6.6%, and there was no difference between arms in opioid-related serious adverse events (hazard ratio 1.03, 95% CI 0.97-1.08) or mortality (hazard ratio 1.00, 95% CI 0.91-1.09) - but patients at accountability-arm facilities were less likely to receive a case review at all (hazard ratio 0.91, 95% CI 0.87-0.95). The implementation tracking study reports the same result at facility level: median completion 71% (interquartile range 48-95%), 18 of 89 surveyed facilities (20%) meeting the 97% target, and facilities given the plain mandate meeting it more often than facilities given the accountability language, 30% against 11% (p=0.04). The practice associated with higher completion was regular self-monitoring and adaptation inside the facility (adjusted incidence rate ratio 1.40) rather than reporting upward from it; dashboard use was reported by 97% of surveyed facilities and local opinion leaders by 80%, while patient-engagement strategies were used by 13%. The oversight-backfire finding belongs to the language experiment alone.
Sources: minegishi2022, strombotne2023, rogal2020
Appears on: /domains/cases/va-storm-opioid-risk
EmpiricalThe benefit and the harm signals from the Veterans Health Administration's mandated opioid case-review policy travel tog…
The benefit and the harm signals from the Veterans Health Administration's mandated opioid case-review policy travel together and neither may be reported alone. The widely quoted mortality result - four-month all-cause mortality odds of 0.78 (95% CI 0.65-0.94) - is an EXPLORATORY endpoint of a trial whose pre-specified primary composite of nine serious-adverse-event categories did not move (odds ratio 0.995, 95% CI 0.875-1.132). Among patients NEWLY DIAGNOSED with opioid use disorder during the trial (28,251 analyzed, estimated off the stepped-wedge threshold-expansion randomization rather than the accountability-language arm) the mandate was associated with 90-day all-cause mortality odds of 1.74 (95% CI 1.06-2.87) with no significant change in serious adverse events; a post-hoc subgroup with an opioid prescription before but not after diagnosis showed 5.87 (95% CI 1.85-18.58). Nearly every quantitative source on this deployment is a department-affiliated research-operations partnership: this is peer-reviewed agency self-evaluation that published its own null, backfire and harm findings, and it is not third-party replication.
Sources: strombotne2023, auty2023
Appears on: /domains/cases/va-storm-opioid-risk