What it is
Monitoring that nobody acts on is only a record of what went wrong. Escalation gives each alert a consequence: more reviewers on the flagged work, a second read of every affected case, or a lower bar for sending a case to a supervisor. The documented cases cited below show why the extra checking helps. Reviewers caught what the model missed because they knew things about the case that the model did not.
What it pushes on in the Lab
In the Lab, this lever raises the checking and correcting that the people using the system can do. It also weakens the pathway by which people adopt the automated system's failed output.
You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.
In the modes that offer aiming, you can aim this lever at particular parts and pathways of a network. Otherwise it applies to the whole network.
Its pattern in the Practice Library
The Practice Library describes the pattern behind this lever:State-feedback vigilance
The pressures it answers
These pressures list this lever among the levers that answer them:
A lever answers a pressure when it pushes the other way on something the pressure pushes on.
Where you can pull it
Networks in the Lab that offer this lever:144
- A commercial code assistant across three enterprises
- A contact centre's generative-AI agent assist
- A heavy-industry predictive-maintenance deployment
- Accelerated Safety Analysis Protocol (ASAP Tool)
- Advance Alert Monitor (AAM) deterioration model
- Air Canada chatbot
- Albert France Services
- Allegheny Family Screening Tool
Every network that offers it
- Allegheny Hello Baby
- Allegheny Housing Assessment
- Amazon fulfillment-centre algorithmic management
- Amazon recruiting engine
- Ambient scribe RCT + monitoring playbook
- Amsterdam Smart Check
- An ambient AI scribe at a multi-specialty health system
- Apple Card underwriting
- Arkansas ARChoices / ARIA
- Audi press-shop inspection
- Automated visual inspection of injectable drugs
- BAMF dialect recognition
- Benefits Data Trust wind-down
- BMW AIQX inspection
- BOSCO (Spain)
- Burokratt
- Caddy adviser copilot at Citizens Advice
- Calgary Drop-In Centre
- CHAI (chronic-homelessness prediction)
- Character.AI crisis-safety stack
- Chicago Public Schools On-Track indicator
- Cigna's PxDx post-service claim review
- Cleveland State remote proctoring
- CNAF benefit-fraud risk score (France)
- CNET AI-drafted articles
- Community Notes on X, formerly Birdwatch on Twitter
- CORA, the DC CFSA policy assistant
- CrimSAFE criminal-record tenant screening
- Crisis Text Line & Loris.ai
- Danske Bank fraud scoring
- Douglas County Decision Aide
- DPD customer-support chatbot
- DWP Whitemail Insights and Vulnerability Scanner
- Eckerd Rapid Safety Feedback
- Enova's CashNetUSA and NetCredit loan servicing
- Epic Sepsis Model
- Equifax's Online Model Server
- EviCore by Evernorth prior-authorization screening
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Fraud false positives that froze real accounts
- Frida (NAV Norway)
- Gaggle Safety Management
- Gated coding-assistant rollout at a regulated bank
- GDS Microsoft 365 Copilot cross-government experiment
- GetCalFresh
- GitHub Copilot at ZoomInfo
- Google ML code completion
- Google's child-safety detection and account enforcement
- GOV.UK Chat
- Hackney / Xantura Early Help Profiling
- Home Office IPIC
- ID.me identity verification
- IDx-DR Autonomous Screening
- Illinois Rapid Safety Feedback
- Imagine LA Benefit Navigator copilot
- Indiana / IBM eligibility modernization
- Insight Bristol / Think Family Database
- INSS automated benefit analysis
- IRS collection chatbots
- Justice Transcribe
- Kaiser Permanente ambient AI scribe
- Kaiser Permanente Suicide-Risk Model
- Klarna AI assistant
- LA County Homelessness Prevention Unit
- LA's coordinated-entry triage revision
- Learned Hand AI clerk pilot (LA and Riverside courts)
- Limbic Access (NHS Talking Therapies)
- London's Strategic Insights Tool
- Los Angeles County Project AURA
- LyssnCrisis counselor QA at ProtoCall Services (988)
- Magic Notes (Beam)
- Meta content enforcement
- Meta employment-ad targeting and delivery optimization
- Meta's cross-check secondary review programme
- Minute / Local Transcribe
- ML anti-money-laundering as primary monitoring
- NarxCare
- Nava assistive benefits chatbot
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- Nevada DETR generative-AI unemployment appeals
- New Zealand MSD Predictive Risk Modelling
- nH Predict Utilization Review
- NYC MyCity business chatbot
- NYC Teenspace
- ODMAP overdose spike alerts
- OPTN eGFR Waiting-Time Correction
- Oregon Safety at Screening
- Oxevision camera monitoring on NHS mental health wards
- Predictive maintenance on a high-speed rail fleet
- Propel in-app SNAP benefits assistant
- REACH VET
- Robodebt (Australia)
- Rotterdam welfare-fraud risk model
- Samagra Vedika
- San Jose's camera car
- Santa Clara County Homelessness Prevention System
- Sepsis Watch deep-learning detection system
- Serbia Social Card (Socijalna karta)
- Singapore's chatbot fleet refresh
- Sistema Alerta Niñez (Chile)
- Sports Illustrated AI bylines
- SSA Insight
- Stratification Tool for Opioid Risk Mitigation
- SyRI (Netherlands)
- Tennessee TennCare TEDS
- Tessa chatbot replacing the NEDA eating-disorder helpline
- The Digit automated-savings tool, or Oportun Set & Save
- The GIFCT hash-sharing database and member matching system
- The NCMEC CyberTipline reporting and triage system
- The Sama Nairobi content-moderation workforce for Meta
- The same AI running hands off: the agentic office
- The same AI with a human checking: the supervised office
- TikTok's EU and UK content-moderation operation
- Trelleborg's Welfare Robot
- TREWS sepsis early-warning system
- Udbetaling Danmark data-driven control (Denmark)
- UK DWP Universal Credit Advances fraud model
- UK Home Office asylum AI copilots
- Unilever and HireVue graduate hiring
- United Behavioral Health's Level of Care Guidelines
- UPS delivery route optimization
- US Birth Match
- VA claims automation (automated survivor-benefit decisions)
- Vanderbilt VSAIL suicide-risk alert
- VI-SPDAT
- Viz.ai LVO Stroke Triage
- What Works for Children's Social Care ML pilots
- Wikipedia's edit-scoring service (ORES, now Lift Wing)
- Wisconsin DEWS
- Woebot (a governed app wind-down)
- Workday AI screening
- Workforce Australia Targeted Compliance Framework
- X Multilingual Hate-Speech Enforcement
- Xantura OneView (predictive homelessness flagging)
- YouTube Covid-19 enforcement
- YouTube's Content ID copyright matching system
The evidence behind its effects
The Lab cites these claims from the evidence registry for this lever's effects.
In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.[4]
Rittenhouse, Algorithms, Humans and Racial Disparities in Child Protective Services https://krittenh.github.io/katherine-rittenhouse.com/Rittenhouse_Algorithms.pdf
https://krittenh.github.io/katherine-rittenhouse.com/Rittenhouse_Algorithms.pdf
Grounds: model org: allegheny_afst
Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf
Appears in: PAN framework development
Grounds: capability governance: at-node control; model org: allegheny_afst
Topics: child-welfare
Centre for Social Data Analytics (AUT), AFST evaluation summary https://csda.aut.ac.nz/news-and-events/2019/allegheny-family-screening-tool-evaluation-improved-decision-accuracy,-reduced-disparities
Grounds: model org: allegheny_afst
Stapleton, L., Lee, M. H., Qing, D., Wright, M., Chouldechova, A., Holstein, K., Wu, Z. S., & Zhu, H. (2022). Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders. 2022 ACM Conference on Fairness Accountability and Transparency, 1162–1177. https://doi.org/10.1145/3531146.3533177
doi.org/10.1145/3531146.3533177
Appears in: PAN framework development; Paramerge authored research
Grounds: deployment audit: Allegheny AFST
Topics: algorithmic-fairness, child-welfare
In the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations erred at about 85% without human review versus 44% with it.[4]
Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
Grounds: model org: michigan_midas
IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Appears in: PAN framework development
Grounds: deployment audit: Michigan MiDAS
AI Incident Database, Incident 373 (MiDAS false fraud claims) https://incidentdatabase.ai/cite/373/
https://incidentdatabase.ai/cite/373/
Grounds: model org: michigan_midas
Benefits Tech Advocacy Hub, Michigan UI False Fraud Determinations https://www.btah.org/case-study/michigan-unemployment-insurance-false-fraud-determinations.html
https://www.btah.org/case-study/michigan-unemployment-insurance-false-fraud-determinations.html
Grounds: model org: michigan_midas
In contextual inquiries with Allegheny AFST call screeners, workers calibrated reliance using contextual case knowledge unavailable to the model and reliably detected and overrode erroneous risk scores — complementary human information, not generic distrust, was the safeguard's mechanism.[2]
Kawakami, A., Sivaraman, V., Cheng, H.-F., Stapleton, L., Cheng, Y., Qing, D., Perer, A., Wu, Z. S., Zhu, H., & Holstein, K. (2022). Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support. In CHI Conference on Human Factors in Computing Systems (CHI '22). ACM. https://doi.org/10.1145/3491102.3517439
doi.org/10.1145/3491102.3517439
Appears in: PAN framework development
Topics: algorithmic-fairness, child-welfare, human-ai-interaction
De-Arteaga, M., Fogliato, R., & Chouldechova, A. (2020). A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores. In CHI Conference on Human Factors in Computing Systems (CHI 2020). ACM. https://doi.org/10.1145/3313831.3376638
doi.org/10.1145/3313831.3376638
Appears in: Evidence reverification (2026)
Topics: algorithmic-fairness, child-welfare, human-ai-interaction