Evidence · The claim ledger
Field benchmarks & evaluations7
Every cited claim this site makes in this evidence area, with the sources that ground it. Source keys link back to the full reference lists on the Evidence Registry.
EmpiricalIn an independent validation — against the NICE Evidence Standards Framework — of a Magic Notes documentation-assistance…
In an independent validation — against the NICE Evidence Standards Framework — of a Magic Notes documentation-assistance pilot at Kent County Council adult social care, staff self-reported weekly written-admin time falling roughly 6.8-7.2 hours (about 35-41%), records submitted some 2.0-3.5 days sooner, and case-note detail rated 6.2 to 8.7 out of 10; the validator judged the findings directionally valid rather than a productivity measurement, because the study was commissioned by the vendor (Beam) — which collected and analysed the data while the validator only sense-checked it — and rested on 29 opt-in staff over 8 weeks with self-estimated time, no control group, no statistical testing, and safety and accuracy explicitly out of scope.
Sources: unityinsights2025, beam, somersetcouncil
Appears on: /pan-lab, /what-ai-can-do
EmpiricalIn the same independent validation of the Kent County Council Magic Notes documentation-assistance pilot, the report rec…
In the same independent validation of the Kent County Council Magic Notes documentation-assistance pilot, the report records the deskilling concern that runs alongside the benefit: one client declined having their session recorded "due to personal feelings of risk of loss of practitioner skills" (p.19). It is a single qualitative observation from a vendor-commissioned pilot of 29 opt-in staff over 8 weeks with no control group — evidence for the direction of the crutch/deskilling risk that accompanies documentation assistance, not for its magnitude.
Sources: unityinsights2025
Appears on: /pan-lab, /what-ai-can-do
EmpiricalIn a peer-reviewed staggered-deployment study of 5,172 customer-support agents at a single firm, access to a generative-…
In a peer-reviewed staggered-deployment study of 5,172 customer-support agents at a single firm, access to a generative-AI assistant raised issues resolved per hour by about 15% on average, with the gain concentrated in the least-experienced workers — roughly +30% for novices versus near-zero for the most experienced, who showed small quality declines; the widely cited 14%/34% pair comes from the 2023 draft, while the peer-reviewed figures are 15%/30%, and because the domain is customer support the direction is imported to social services but the magnitude is never treated as a fixed quantity.
Sources: brynjolfsson2025a
Appears on: /pan-lab, /what-ai-can-do
EmpiricalIn a randomized controlled trial of a benefits-navigation chatbot (co-authored by Cornell researchers and the tool's dev…
In a randomized controlled trial of a benefits-navigation chatbot (co-authored by Cornell researchers and the tool's developer, Nava) with 125 caseworkers across six Los Angeles County organizations over 14 weeks, caseworkers answered complex benefit questions at about 49% accuracy unaided, and high-quality chatbot suggestions raised accuracy by roughly 27 percentage points — with larger gains on harder questions, but a persistent 'AI underreliance' plateau in which correct suggestions were not always adopted; the trial did not establish a clear effect on administrative burden, a null reported honestly rather than inferred as a benefit.
Sources: gosciak2026, kanne2025, navapublicbenefitcorporation2025d, navapublicbenefitcorporation2026
Appears on: /pan-lab, /what-ai-can-do
EmpiricalIn the county-commissioned impact evaluation of the Allegheny Family Screening Tool, screen-in accuracy — further action…
In the county-commissioned impact evaluation of the Allegheny Family Screening Tool, screen-in accuracy — further action or re-referral within 60 days — rose from 42.85% to 46.61% (p=.000) while consistency across screeners was maintained; the finding is contested and quasi-experimental, with true maltreatment rates unknown and the accuracy gains concentrated among white children and ages 7-12, the gain for Black children attenuating to statistical non-significance.
Sources: goldhaberfiebertprince2019
Appears on: /pan-lab, /what-ai-can-do
EmpiricalCrisis-text triage classifiers read language, so the language they read least well is where they fail. The volume's ment…
Crisis-text triage classifiers read language, so the language they read least well is where they fail. The volume's mental-health chapter reports that sarcasm, cultural idiom and code-switching confuse these models, that benchmark studies show markedly lower accuracy on African American English and other underrepresented dialects, and that the operational cost runs in both directions: a false positive can trigger an unwanted welfare check, a false negative leaves a texter waiting in silence. This is benchmark evidence about a class of classifier, not a measurement of any deployment in this registry. The underlying benchmark study is not held in this repository's reference snapshot, so no magnitude is carried, and the finding must never be attached to any named service's own published or unpublished figures. Nothing here is a fairness metric and nothing here is computed about any texter.
Sources: yang2026c
Appears on: /domains/cases/crisis-text-line-loris, /pan-lab
EmpiricalThe volume's substance-use chapter reports the field's one sustained deployment benefit case: a national health system's…
The volume's substance-use chapter reports the field's one sustained deployment benefit case: a national health system's opioid risk-mitigation dashboard, an advisory clinical decision-support tool that stratifies patients for review, was evaluated in a randomized design across the system's medical centers, and use of the risk stratification was reported as associated with a decrease in mortality among the at-risk patients it covered. Three limits travel with the finding and are part of the claim. It is an association reported in a peer-reviewed secondary synthesis, not a causal result this repository can inspect: the primary evaluation is absent from the reference snapshot, so no sample size, facility count, follow-up window or effect size is carried here. It is a benefit direction on a domain this registry otherwise describes almost entirely in the failure register, and it is recorded so that the failure register is not mistaken for the whole evidence base. And it is an outcome for people who are outside the model by construction: it is never read from a Lab gauge, never a Service-Regime number, and never computed from any diagram.
Sources: saba2026
Appears on: /domains/behavioral-health-triage