ParamergeParamerge

Evidence · The claim ledger

Model reliability & evaluation limits11

Every cited claim this site makes in this evidence area, with the sources that ground it. Source keys link back to the full reference lists on the Evidence Registry.

EmpiricalProfessional-verification cultures documented in social work practice sustain peer checking of AI output rather than unq…

Professional-verification cultures documented in social work practice sustain peer checking of AI output rather than unquestioned acceptance.

Sources: baez2026

Appears on: /pan-lab

EmpiricalRetrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain unfaithful…

Retrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain unfaithful even when the retrieved passage is correct, so faithfulness is bounded rather than assured.

Sources: faithfulrag, faithfulragwithsparseautoenc, ragevaluationsurvey

Appears on: /pan-lab, /practice/curated-corpus-retrieval

EmpiricalRestricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented systems fac…

Restricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented systems fact-checking against a curated peer-reviewed corpus reach roughly 0.97+ accuracy and factuality evaluation is limited by knowledge-base coverage — what is checkable depends on what is documented — so a vetted corpus reduces contamination drawn back into the model relative to open retrieval, though faithfulness remains imperfect under knowledge conflict.

Sources: retrievalaugmentedcovidfactc, ragevaluationsurvey, faithfulrag

Appears on: /pan-lab, /practice/curated-corpus-retrieval

EmpiricalAutomated output checks are partial, not complete — measured detector-accuracy bands sit well below completeness, especi…

Automated output checks are partial, not complete — measured detector-accuracy bands sit well below completeness, especially on hard or adversarial content.

Sources: theillusionofprogress, halogen, datadogllmasajudge2025, mentalhealthchatbotdetection

Appears on: /pan-lab, /practice/bounded-output-screening

EmpiricalModel error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly …

Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.

Sources: xuetal2024, karpowicz2025, halogen, openai2025, llmstats2026, suprmindbenchmarkdigest2026

Appears on: /pan-lab, /practice/improve-the-model

EmpiricalAutomated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings versus about …

Automated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings versus about 55% on hard content and 9.3% recall in worst-case measurements.

Sources: faithfulragleaderboard, theillusionofprogress, mentalhealthchatbotdetection, datadogllmasajudge2025, samedetectionaccuracyliterat

Appears on: /pan-lab, /practice/bounded-output-screening

EmpiricalRecord audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy falls away on…

Record audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy falls away on hard content, so decontamination is bounded rather than total.

Sources: halludetectlegaldomain, samedetectionaccuracyliterat, theillusionofprogress

Appears on: /pan-lab, /practice/content-aware-decontamination

EmpiricalThe verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliability and co…

The verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliability and collapses under knowledge conflict.

Sources: retrievalaugmentedcovidfactc, faithfulrag, faithfulragwithsparseautoenc

Appears on: /pan-lab, /practice/curated-corpus-retrieval

AssumptionThe verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibration-requ…

The verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibration-required and has never been measured.

Sources: ragevaluationsurvey

Appears on: /pan-lab

EmpiricalLanguage models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consist…

Language models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consistent with it — recognizing 67-87% of those fabrications as false when re-asked in a clean, uncontaminated context but not correcting them in place — so one error deterministically spawns supporting errors, a self-sustaining failure the model's own downstream output feeds.

Sources: zhang2024

Appears on: /pan-lab

EmpiricalThe advocacy chapter's systematic review screened 7,715 records, included 415 articles and described 80 of them in detai…

The advocacy chapter's systematic review screened 7,715 records, included 415 articles and described 80 of them in detail, reporting cluster sizes across five method-application buckets. Those are counts of what has been built and published: the review reports no pooled effectiveness estimate and no risk-of-bias appraisal, and its own limitations section calls for randomised trials or rigorous observational studies to become standard practice. A count of applications measures activity, never effect.

Sources: alba2026

Appears on: /pan-lab