ParamergeParamerge

Practice Library

Governance patternstructural

Improve the model

The default instinct — buy or build a better model — is a real lever with an honest, limited reach.

What it changes

decreasedAutonomous systemSociotechnical simulation result: ≈6% of the harm removed by a model upgrade alone in the same sociotechnical simulation that credited a verifier with ≈46%.

Who can pull it

DeveloperVendorDeploying organization

What it looks like institutionally

Reducing the model's base error rate helps every downstream pathway a little. It is also the lever institutions reach for first, because it requires no organizational change: procurement instead of governance.

Its honest limits: base error has an empirical floor (no current system reaches zero in demanding domains), and in propagation terms a better model shrinks the source while leaving every loop — adoption, records, retrieval, peer spread — untouched. Sociotechnical simulation runs across hundreds of deployment structures found system-side levers outperforming equal-effort model improvements in the overwhelming majority of cases; the ledgered scenario results carry the specifics and their caveats.

Use it, but use it last-alone: pair model improvements with the structural levers that govern what happens to the errors that remain.

Ledgered PAN-run results used above

In the sociotechnical simulation, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.[]

In the sociotechnical simulation, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.[]

Addresses: High base error. Test a version of this lever in the PAN Lab.

Deciding whether this lever fits your deployment?

Which patterns matter, and in what order, depends on your system's actual shape. Ranking your options on evidence, with what can backfire stated, is engagement work.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

ScenarioIn the sociotechnical simulation, over a supervised-plus-agent scenario, adding a verifier to the autonomous a…

In the sociotechnical simulation, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.

From the sociotechnical simulation: PAN social-work governance guidance, lever-ranking comparison.

ScenarioIn the sociotechnical simulation, fixing the surrounding system out-leveraged an equal-effort model upgrade in…

In the sociotechnical simulation, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.

From the sociotechnical simulation: PAN baseline analysis.

EmpiricalModel error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors ru…

Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.

xuetal2024GroundingPreprintSave

Xu et al. (2024), 'Hallucination is Inevitable: An Innate Limitation of LLMs' — formal proof that hallucination cannot be eliminated.

Grounds: empirical cap: model_error_base (min)

karpowicz2025GroundingPreprintSave

Karpowicz (2025) — three independent mathematical frameworks (auction theory, proper scoring, log-sum-exp) all conclude no LLM inference mechanism can be simultaneously truthful, etc.

Grounds: empirical cap: model_error_base (min)

halogenGroundingPeer-reviewedSave

HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292

https://arxiv.org/abs/2501.08292

Grounds: empirical cap: model_error_base (min)

openai2025GroundingFrontier labSave

OpenAI (2025), 'Why Language Models Hallucinate' — next-token training plus IDK-penalizing benchmarks push models to bluff; explains the persistent nonzero floor.

Grounds: empirical cap: model_error_base (min)

llmstats2026GroundingIndustry evaluationSave

llm-stats.com failure-focused eval (2026) — FactsGrounding 89.1% accuracy => ~10.9% failure on a relatively easy grounded benchmark.

Grounds: empirical cap: model_error_base (min)

suprmindbenchmarkdigest2026GroundingIndustry evaluationSave

Suprmind benchmark digest (2026) — production ChatGPT ~4.8% major-incorrect with reasoning vs ~11.6% without; HealthBench 3.6%->1.6% with GPT-5 thinking.

Grounds: empirical cap: model_error_base (min)