Evidence · The claim ledger
Model reliability & evaluation limits11
Every cited claim this site makes in this evidence area, with the sources that ground it. Source keys link back to the full reference lists on the Evidence Registry.
EmpiricalProfessional-verification cultures documented in social work practice sustain peer checking of AI output rather than unq…
Professional-verification cultures documented in social work practice sustain peer checking of AI output rather than unquestioned acceptance.
Sources: baez2026
Appears on: /pan-lab
EmpiricalRetrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain unfaithful…
Retrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain unfaithful even when the retrieved passage is correct, so faithfulness is bounded rather than assured.
Sources: faithfulrag, faithfulragwithsparseautoenc, ragevaluationsurvey
Appears on: /pan-lab, /practice/curated-corpus-retrieval
EmpiricalRestricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented systems fac…
Restricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented systems fact-checking against a curated peer-reviewed corpus reach roughly 0.97+ accuracy and factuality evaluation is limited by knowledge-base coverage — what is checkable depends on what is documented — so a vetted corpus reduces contamination drawn back into the model relative to open retrieval, though faithfulness remains imperfect under knowledge conflict.
Sources: retrievalaugmentedcovidfactc, ragevaluationsurvey, faithfulrag
Appears on: /pan-lab, /practice/curated-corpus-retrieval
EmpiricalAutomated output checks are partial, not complete — measured detector-accuracy bands sit well below completeness, especi…
Automated output checks are partial, not complete — measured detector-accuracy bands sit well below completeness, especially on hard or adversarial content.
Sources: theillusionofprogress, halogen, datadogllmasajudge2025, mentalhealthchatbotdetection
Appears on: /pan-lab, /practice/bounded-output-screening
EmpiricalModel error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly …
Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.
Sources: xuetal2024, karpowicz2025, halogen, openai2025, llmstats2026, suprmindbenchmarkdigest2026
Appears on: /pan-lab, /practice/improve-the-model
EmpiricalAutomated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings versus about …
Automated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings versus about 55% on hard content and 9.3% recall in worst-case measurements.
Sources: faithfulragleaderboard, theillusionofprogress, mentalhealthchatbotdetection, datadogllmasajudge2025, samedetectionaccuracyliterat
Appears on: /pan-lab, /practice/bounded-output-screening
EmpiricalRecord audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy falls away on…
Record audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy falls away on hard content, so decontamination is bounded rather than total.
Sources: halludetectlegaldomain, samedetectionaccuracyliterat, theillusionofprogress
Appears on: /pan-lab, /practice/content-aware-decontamination
EmpiricalThe verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliability and co…
The verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliability and collapses under knowledge conflict.
Sources: retrievalaugmentedcovidfactc, faithfulrag, faithfulragwithsparseautoenc
Appears on: /pan-lab, /practice/curated-corpus-retrieval
AssumptionThe verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibration-requ…
The verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibration-required and has never been measured.
Sources: ragevaluationsurvey
Appears on: /pan-lab
EmpiricalLanguage models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consist…
Language models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consistent with it — recognizing 67-87% of those fabrications as false when re-asked in a clean, uncontaminated context but not correcting them in place — so one error deterministically spawns supporting errors, a self-sustaining failure the model's own downstream output feeds.
Sources: zhang2024
Appears on: /pan-lab
EmpiricalThe advocacy chapter's systematic review screened 7,715 records, included 415 articles and described 80 of them in detai…
The advocacy chapter's systematic review screened 7,715 records, included 415 articles and described 80 of them in detail, reporting cluster sizes across five method-application buckets. Those are counts of what has been built and published: the review reports no pooled effectiveness estimate and no risk-of-bias appraisal, and its own limitations section calls for randomised trials or rigorous observational studies to become standard practice. A count of applications measures activity, never effect.
Sources: alba2026
Appears on: /pan-lab