What it is
An upgrade to the model is the fix most organizations reach for first. The vendor retrains the model, the team swaps in a newer version, or someone tunes it on local cases. The formal results cited below rule out a model that makes no errors at all. Even the best upgrade leaves errors for the rest of the system to catch.
What it pushes on in the Lab
In the Lab, this lever lowers the errors the automated system produces. It changes no pathway by which an error enters people's work or the records. The remaining errors do as much harm as those pathways allow.
You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.
In the modes that offer aiming, you can aim this lever at particular parts and pathways of a network. Otherwise it applies to the whole network.
Its pattern in the Practice Library
The Practice Library describes the pattern behind this lever:Improve the model
The pressures it answers
These pressures list this lever among the levers that answer them:
A lever answers a pressure when it pushes the other way on something the pressure pushes on.
Where you can pull it
Networks in the Lab that offer this lever:168
- A commercial code assistant across three enterprises
- A contact centre's generative-AI agent assist
- A heavy-industry predictive-maintenance deployment
- Accelerated Safety Analysis Protocol (ASAP Tool)
- Advance Alert Monitor (AAM) deterioration model
- Air Canada chatbot
- Albert France Services
- Allegheny Family Screening Tool
Every network that offers it
- Allegheny Hello Baby
- Allegheny Housing Assessment
- Amazon Flex driver standing and deactivation
- Amazon fulfillment-centre algorithmic management
- Amazon fulfillment-centre productivity discipline
- Amazon recruiting engine
- Ambient scribe RCT + monitoring playbook
- Amsterdam Smart Check
- An ambient AI scribe at a multi-specialty health system
- Aon's three-instrument pre-hire assessment suite
- Apple Card underwriting
- Arkansas ARChoices / ARIA
- Audi press-shop inspection
- Automated visual inspection of injectable drugs
- BAMF dialect recognition
- Benefits Data Trust wind-down
- BMW AIQX inspection
- BOSCO (Spain)
- Burokratt
- CA-CDS Child Abuse Alerting
- Caddy adviser copilot at Citizens Advice
- Calgary Drop-In Centre
- CDTFA Axyom Assist
- CHAI (chronic-homelessness prediction)
- Character.AI crisis-safety stack
- Checkr's automated background-check platform
- Cigna's PxDx post-service claim review
- Cleveland State remote proctoring
- CNAF benefit-fraud risk score (France)
- CNET AI-drafted articles
- Colorado Family Safety and Risk Assessments
- Community Notes on X, formerly Birdwatch on Twitter
- Cost-Proxy Care Stratification
- Credit Acceptance's net-collections Score, inside CAPS
- Dave ExtraCash (CashAI)
- Douglas County Decision Aide
- DWP Whitemail Insights and Vulnerability Scanner
- Earnest AI underwriting
- Eckerd Rapid Safety Feedback
- EDD Virtual Assistant
- Enova's CashNetUSA and NetCredit loan servicing
- Equifax's Online Model Server
- EviCore by Evernorth prior-authorization screening
- Family-Match (Adoption-Share)
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Fraud false positives that froze real accounts
- Frida (NAV Norway)
- Gaggle Safety Management
- GDS Microsoft 365 Copilot cross-government experiment
- GetCalFresh
- Gladsaxe model
- Google ML code completion
- GOV.UK Chat
- Hackney / Xantura Early Help Profiling
- HireVue's video interview and assessment platform
- Home Office IPIC
- Homebase Risk Assessment Questionnaire
- IBM Watson for Oncology
- ID.me identity verification
- IDx-DR Autonomous Screening
- Illinois DCFS Augintel
- Illinois Rapid Safety Feedback
- Imagine LA Benefit Navigator copilot
- Indiana / IBM eligibility modernization
- Insight Bristol / Think Family Database
- INSS automated benefit analysis
- Intuit's recorded video assessment for promotion
- IRS collection chatbots
- Justice Transcribe
- Kaiser Permanente ambient AI scribe
- Kaiser Permanente Suicide-Risk Model
- Klarna AI assistant
- LA County Homelessness Prevention Unit
- LA's coordinated-entry triage revision
- Learned Hand AI clerk pilot (LA and Riverside courts)
- Limbic Access (NHS Talking Therapies)
- London's Strategic Insights Tool
- Los Angeles County Project AURA
- LyssnCrisis counselor QA at ProtoCall Services (988)
- M-Shwari & Kenya's Digital Credit Market
- Magic Notes (Beam)
- Mass.gov Virtual Assistant
- McHire, McDonald's franchise hiring platform
- Medicaid unwinding ex-parte renewals
- Meta content enforcement
- Meta's cross-check secondary review programme
- Michigan MiDAS
- Michigan MiDAS
- Minute / Local Transcribe
- MyFriendBen benefits screener
- NarxCare
- Navy Federal mortgage underwriting
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- Nevada DETR generative-AI unemployment appeals
- New Zealand MSD Predictive Risk Modelling
- nH Predict Utilization Review
- NJ AI Assistant
- NYC MyCity business chatbot
- ODMAP overdose spike alerts
- OPTN eGFR Waiting-Time Correction
- Oregon Safety at Screening
- Oxevision camera monitoring on NHS mental health wards
- Practice Fusion Pain CDS
- Predict-Align-Prevent
- Predictive maintenance on a high-speed rail fleet
- ProKid (Netherlands)
- pymetrics Soft Skills Platform cooperative audit
- REACH VET
- Robodebt (Australia)
- Robodebt (Australia)
- Rotterdam welfare-fraud risk model
- SafeRent Tenant Screening Score
- Samagra Vedika
- San Jose's camera car
- Santa Clara County Homelessness Prevention System
- Santander Consumer USA's loss forecasting score
- Sepsis Watch deep-learning detection system
- Serbia Social Card (Socijalna karta)
- Singapore's chatbot fleet refresh
- Sirius XM Radio's iCIMS-based applicant screening
- Sistema Alerta Niñez (Chile)
- SSA 800-Number Conversational AI Assistant
- SSA Insight
- StopNCII & Take It Down
- Stratification Tool for Opioid Risk Mitigation
- SyRI (Netherlands)
- Tennessee TennCare TEDS
- Tessa chatbot replacing the NEDA eating-disorder helpline
- The Digit automated-savings tool, or Oportun Set & Save
- The NCMEC CyberTipline reporting and triage system
- The same AI running hands off: the agentic office
- The same AI under full guardrails: the professional office
- The same AI with a human checking: the supervised office
- TikTok's EU and UK content-moderation operation
- TransUnion OFAC Name Screen
- Trelleborg's Welfare Robot
- TREWS sepsis early-warning system
- Udbetaling Danmark data-driven control (Denmark)
- UK DWP Universal Credit Advances fraud model
- UK Home Office asylum AI copilots
- Unilever and HireVue graduate hiring
- United Behavioral Health's Level of Care Guidelines
- UPS delivery route optimization
- Upstart lending model
- US Birth Match
- VA claims automation (automated survivor-benefit decisions)
- Vanderbilt VSAIL suicide-risk alert
- VI-SPDAT
- Viz.ai LVO Stroke Triage
- Wells Fargo refinance underwriting (CORE/ECS)
- What Works for Children's Social Care ML pilots
- Wikipedia's edit-scoring service (ORES, now Lift Wing)
- Wisconsin DEWS
- Woebot (a governed app wind-down)
- Workday AI screening
- Workforce Australia Targeted Compliance Framework
- X Multilingual Hate-Speech Enforcement
- Xantura OneView (predictive homelessness flagging)
- YouTube Covid-19 enforcement
- YouTube's Content ID copyright matching system
The evidence behind its effects
The Lab cites these claims from the evidence registry for this lever's effects.
In the sociotechnical simulation, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.[sim]
In the sociotechnical simulation, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.[sim]
Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.[6]
Xu et al. (2024), 'Hallucination is Inevitable: An Innate Limitation of LLMs' — formal proof that hallucination cannot be eliminated.
Grounds: empirical cap: model_error_base (min)
Karpowicz (2025) — three independent mathematical frameworks (auction theory, proper scoring, log-sum-exp) all conclude no LLM inference mechanism can be simultaneously truthful, etc.
Grounds: empirical cap: model_error_base (min)
HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292
https://arxiv.org/abs/2501.08292
Grounds: empirical cap: model_error_base (min)
OpenAI (2025), 'Why Language Models Hallucinate' — next-token training plus IDK-penalizing benchmarks push models to bluff; explains the persistent nonzero floor.
Grounds: empirical cap: model_error_base (min)
llm-stats.com failure-focused eval (2026) — FactsGrounding 89.1% accuracy => ~10.9% failure on a relatively easy grounded benchmark.
Grounds: empirical cap: model_error_base (min)
Suprmind benchmark digest (2026) — production ChatGPT ~4.8% major-incorrect with reasoning vs ~11.6% without; HealthBench 3.6%->1.6% with GPT-5 thinking.
Grounds: empirical cap: model_error_base (min)
How a model scores on data held back from its own training and how it scores at a different site are different quantities. The volume's disability chapter reports a named pair where the second is materially lower than the first. A figure quoted without saying which of the two it is does not tell a reader what the model will do in their setting.[†]
Wang, J., & Begg, M. D. (2026). AI for Physical, Cognitive, and Developmental Challenges. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_6
doi.org/10.1007/978-3-032-18443-6_6
Appears in: AI in Social Work (Springer, 2026)
Topics: disability, human-ai-interaction, social-work
The volume's substance-use chapter reports that nearly all models in the field it reviews are developed and evaluated on a single dataset. Where that holds, a reported ceiling describes that sample rather than a portable capability, and it should be read as the best case observed on one collection of records, not as what the tool will do elsewhere.[†]
Saba, S., & Leibowitz, G. (2026). AI in Substance Use and Addiction Prevention. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_11
doi.org/10.1007/978-3-032-18443-6_11
Appears in: AI in Social Work (Springer, 2026)
Grounds: model org: odmap_overdose_spike_alerts; model org: woebot_health_app
Topics: social-work, substance-use
Discrimination statistics such as AUROC, precision, F1, and lift describe how a model separates cases on a labelled sample. They are not per-interaction rates at which an error is adopted, written into a record, or corrected. The volume's research chapter treats these as different quantities, and they must never be entered into a diagram as though one substitutes for the other.[†]
Yang, Y., Huang, J., & An, R. (2026). AI in Social Work Research. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_21
doi.org/10.1007/978-3-032-18443-6_21
Appears in: AI in Social Work (Springer, 2026)
Topics: research-methods, social-work, social-work-research