What it is
Two copies of one model make the same mistakes at the same time. In the Michigan MiDAS case cited below, one uniform rule set produced tens of thousands of correlated wrongful fraud determinations, one flaw repeated at the scale of the caseload. A second model from another vendor or another design disagrees where the first is wrong in a way the second is not. The research cited below also finds that a model rarely corrects its own answer once it has committed to it. The check therefore has to come from outside the model.
What it pushes on in the Lab
In the Lab, this lever strengthens the pathway by which models cross-check each other's output. It also weakens the pathway by which AI models or agents relay failures to one another. Its side effect raises operator deference drift, because two systems that agree look like proof.
You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.
In the modes that offer aiming, you can aim this lever at particular parts and pathways of a network. Otherwise it applies to the whole network.
Its pattern in the Practice Library
The Practice Library describes the pattern behind this lever:Cross-model verification
The pressures it answers
These pressures list this lever among the levers that answer them:
A lever answers a pressure when it pushes the other way on something the pressure pushes on.
Where you can pull it
Networks in the Lab that offer this lever:112
- A commercial code assistant across three enterprises
- A contact centre's generative-AI agent assist
- A heavy-industry predictive-maintenance deployment
- Accelerated Safety Analysis Protocol (ASAP Tool)
- Advance Alert Monitor (AAM) deterioration model
- Air Canada chatbot
- Amazon Flex driver standing and deactivation
- Amazon recruiting engine
Every network that offers it
- Ambient scribe RCT + monitoring playbook
- Amsterdam Smart Check
- An ambient AI scribe at a multi-specialty health system
- Aon's three-instrument pre-hire assessment suite
- Arkansas ARChoices / ARIA
- Audi press-shop inspection
- Automated visual inspection of injectable drugs
- BAMF dialect recognition
- BMW AIQX inspection
- BOSCO (Spain)
- Burokratt
- CDTFA Axyom Assist
- Character.AI crisis-safety stack
- Checkr's automated background-check platform
- Cigna's PxDx post-service claim review
- Cleveland State remote proctoring
- CNAF benefit-fraud risk score (France)
- Colorado Family Safety and Risk Assessments
- Community Notes on X, formerly Birdwatch on Twitter
- Cost-Proxy Care Stratification
- Danske Bank fraud scoring
- Dave ExtraCash (CashAI)
- DPD customer-support chatbot
- DWP Whitemail Insights and Vulnerability Scanner
- Earnest AI underwriting
- Eckerd Rapid Safety Feedback
- Enova's CashNetUSA and NetCredit loan servicing
- Epic Sepsis Model
- Equifax's Online Model Server
- EviCore by Evernorth prior-authorization screening
- Family-Match (Adoption-Share)
- Fraud false positives that froze real accounts
- Frida (NAV Norway)
- Gated coding-assistant rollout at a regulated bank
- GDS Microsoft 365 Copilot cross-government experiment
- GitHub Copilot at ZoomInfo
- Google ML code completion
- Google's child-safety detection and account enforcement
- GOV.UK Chat
- HireVue's video interview and assessment platform
- Illinois DCFS Augintel
- Illinois Rapid Safety Feedback
- INSS automated benefit analysis
- Intuit's recorded video assessment for promotion
- Justice Transcribe
- Kaiser Permanente ambient AI scribe
- Kaiser Permanente Suicide-Risk Model
- Klarna AI assistant
- LA's coordinated-entry triage revision
- Learned Hand AI clerk pilot (LA and Riverside courts)
- Limbic Access (NHS Talking Therapies)
- Los Angeles County Project AURA
- Mass.gov Virtual Assistant
- McHire, McDonald's franchise hiring platform
- Medicaid unwinding ex-parte renewals
- Meta content enforcement
- Meta employment-ad targeting and delivery optimization
- Meta's cross-check secondary review programme
- Michigan MiDAS
- Minute / Local Transcribe
- ML anti-money-laundering as primary monitoring
- MyFriendBen benefits screener
- NarxCare
- Navy Federal mortgage underwriting
- Nevada DETR generative-AI unemployment appeals
- nH Predict Utilization Review
- NJ AI Assistant
- NYC MyCity business chatbot
- ODMAP overdose spike alerts
- OPTN eGFR Waiting-Time Correction
- Predict-Align-Prevent
- Predictive maintenance on a high-speed rail fleet
- pymetrics Soft Skills Platform cooperative audit
- REACH VET
- Robodebt (Australia)
- Samagra Vedika
- San Jose's camera car
- Santander Consumer USA's loss forecasting score
- Sepsis Watch deep-learning detection system
- Singapore's chatbot fleet refresh
- Sirius XM Radio's iCIMS-based applicant screening
- SSA 800-Number Conversational AI Assistant
- StopNCII & Take It Down
- Stratification Tool for Opioid Risk Mitigation
- Tessa chatbot replacing the NEDA eating-disorder helpline
- The Digit automated-savings tool, or Oportun Set & Save
- The GIFCT hash-sharing database and member matching system
- The same AI running hands off: the agentic office
- TransUnion OFAC Name Screen
- TREWS sepsis early-warning system
- Udbetaling Danmark data-driven control (Denmark)
- UK DWP Universal Credit Advances fraud model
- UK Home Office asylum AI copilots
- Unilever and HireVue graduate hiring
- United Behavioral Health's Level of Care Guidelines
- Upstart lending model
- VI-SPDAT
- Wells Fargo refinance underwriting (CORE/ECS)
- Wikipedia's edit-scoring service (ORES, now Lift Wing)
- Wisconsin DEWS
- Workday AI screening
- Workforce Australia Targeted Compliance Framework
- X Multilingual Hate-Speech Enforcement
- YouTube Covid-19 enforcement
The evidence behind its effects
The Lab cites these claims from the evidence registry for this lever's effects.
Language models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consistent with it — recognizing 67-87% of those fabrications as false when re-asked in a clean, uncontaminated context but not correcting them in place — so one error deterministically spawns supporting errors, a self-sustaining failure the model's own downstream output feeds.[†]
Zhang, M., Press, O., Merrill, W., Liu, A., & Smith, N. A. (2024). How Language Model Hallucinations Can Snowball. In International Conference on Machine Learning (ICML 2024), PMLR 235:59670-59684. https://doi.org/10.48550/arXiv.2305.13534
doi.org/10.48550/arXiv.2305.13534
Appears in: PAN framework development
Topics: ai-safety