Skip to content

PAN Lab levers

Lever

Upgrade model

The organization upgrades the model so that it produces fewer errors. The work is necessary, and it is rarely enough. In the simulation runs cited below, fixing the system around the model did more good than an equal effort spent on the model, in nearly every case tested.

What it is

An upgrade to the model is the fix most organizations reach for first. The vendor retrains the model, the team swaps in a newer version, or someone tunes it on local cases. The formal results cited below rule out a model that makes no errors at all. Even the best upgrade leaves errors for the rest of the system to catch.

What it pushes on in the Lab

In the Lab, this lever lowers the errors the automated system produces. It changes no pathway by which an error enters people's work or the records. The remaining errors do as much harm as those pathways allow.

decreasedAutomated system

You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.

In the modes that offer aiming, you can aim this lever at particular parts and pathways of a network. Otherwise it applies to the whole network.

Its pattern in the Practice Library

The Practice Library describes the pattern behind this lever:Improve the model

The pressures it answers

These pressures list this lever among the levers that answer them:

A lever answers a pressure when it pushes the other way on something the pressure pushes on.

Where you can pull it

Networks in the Lab that offer this lever:168

Every network that offers it

The evidence behind its effects

The Lab cites these claims from the evidence registry for this lever's effects.

In the sociotechnical simulation, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.[sim]

In the sociotechnical simulation, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.[sim]

Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.[6]

xuetal2024GroundingPreprintSave

Xu et al. (2024), 'Hallucination is Inevitable: An Innate Limitation of LLMs' — formal proof that hallucination cannot be eliminated.

Grounds: empirical cap: model_error_base (min)

karpowicz2025GroundingPreprintSave

Karpowicz (2025) — three independent mathematical frameworks (auction theory, proper scoring, log-sum-exp) all conclude no LLM inference mechanism can be simultaneously truthful, etc.

Grounds: empirical cap: model_error_base (min)

halogenGroundingPeer-reviewedSave

HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292

https://arxiv.org/abs/2501.08292

Grounds: empirical cap: model_error_base (min)

openai2025GroundingFrontier labSave

OpenAI (2025), 'Why Language Models Hallucinate' — next-token training plus IDK-penalizing benchmarks push models to bluff; explains the persistent nonzero floor.

Grounds: empirical cap: model_error_base (min)

llmstats2026GroundingIndustry evaluationSave

llm-stats.com failure-focused eval (2026) — FactsGrounding 89.1% accuracy => ~10.9% failure on a relatively easy grounded benchmark.

Grounds: empirical cap: model_error_base (min)

suprmindbenchmarkdigest2026GroundingIndustry evaluationSave

Suprmind benchmark digest (2026) — production ChatGPT ~4.8% major-incorrect with reasoning vs ~11.6% without; HealthBench 3.6%->1.6% with GPT-5 thinking.

Grounds: empirical cap: model_error_base (min)

How a model scores on data held back from its own training and how it scores at a different site are different quantities. The volume's disability chapter reports a named pair where the second is materially lower than the first. A figure quoted without saying which of the two it is does not tell a reader what the model will do in their setting.[†]

wang2026aAcademicSave

Wang, J., & Begg, M. D. (2026). AI for Physical, Cognitive, and Developmental Challenges. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_6

doi.org/10.1007/978-3-032-18443-6_6

Appears in: AI in Social Work (Springer, 2026)

Topics: disability, human-ai-interaction, social-work

The volume's substance-use chapter reports that nearly all models in the field it reviews are developed and evaluated on a single dataset. Where that holds, a reported ceiling describes that sample rather than a portable capability, and it should be read as the best case observed on one collection of records, not as what the tool will do elsewhere.[†]

saba2026AcademicAcademicSave

Saba, S., & Leibowitz, G. (2026). AI in Substance Use and Addiction Prevention. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_11

doi.org/10.1007/978-3-032-18443-6_11

Appears in: AI in Social Work (Springer, 2026)

Grounds: model org: odmap_overdose_spike_alerts; model org: woebot_health_app

Topics: social-work, substance-use

Discrimination statistics such as AUROC, precision, F1, and lift describe how a model separates cases on a labelled sample. They are not per-interaction rates at which an error is adopted, written into a record, or corrected. The volume's research chapter treats these as different quantities, and they must never be entered into a diagram as though one substitutes for the other.[†]

yang2026eAcademicSave

Yang, Y., Huang, J., & An, R. (2026). AI in Social Work Research. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_21

doi.org/10.1007/978-3-032-18443-6_21

Appears in: AI in Social Work (Springer, 2026)

Topics: research-methods, social-work, social-work-research

Pull this lever in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab