What it is
Controls tuned at launch fit the system that launched. A model drifts, a vendor updates it, and the caseload changes, while the alert levels stay where they were set. A review cadence commits the organization to looking again on a schedule. The monitoring research cited below found that some systems behaved more safely when they believed they were watched, so reviews at predictable times are easier to game than reviews at unpredictable ones.
What it pushes on in the Lab
In the Lab, this lever weakens the failure regime, the network's tendency toward failures that compound rather than correct themselves. At its strong tier it also engages deployment authority, and it puts a ceiling on the failure regime.
You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.
The strong tier also pushes on these:
The Lab applies this lever to the whole network. It acts only on gauges that read the whole deployment, so it has no single part to aim at.
Its pattern in the Practice Library
The Practice Library describes the pattern behind this lever:Oversight cadence & retrospectives
The pressures it answers
These pressures list this lever among the levers that answer them:
- Monitoring goes stale
- Agent links sprawlat its strong tier
- Unsanctioned AI useat its strong tier
- Expectations outrun the gains
A lever answers a pressure when it pushes the other way on something the pressure pushes on.
Where you can pull it
Networks in the Lab that offer this lever:185
- A commercial code assistant across three enterprises
- A contact centre's generative-AI agent assist
- A heavy-industry predictive-maintenance deployment
- Accelerated Safety Analysis Protocol (ASAP Tool)
- Advance Alert Monitor (AAM) deterioration model
- Air Canada chatbot
- Albert France Services
- Allegheny Family Screening Tool
Every network that offers it
- Allegheny Hello Baby
- Allegheny Housing Assessment
- Amazon Flex driver standing and deactivation
- Amazon fulfillment-centre algorithmic management
- Amazon fulfillment-centre productivity discipline
- Amazon recruiting engine
- Ambient scribe RCT + monitoring playbook
- Amsterdam Smart Check
- An ambient AI scribe at a multi-specialty health system
- Aon's three-instrument pre-hire assessment suite
- Apple Card underwriting
- Arkansas ARChoices / ARIA
- Audi press-shop inspection
- Automated visual inspection of injectable drugs
- BAMF dialect recognition
- Benefits Data Trust wind-down
- BMW AIQX inspection
- BOSCO (Spain)
- Burokratt
- CA-CDS Child Abuse Alerting
- Caddy adviser copilot at Citizens Advice
- Calgary Drop-In Centre
- CDTFA Axyom Assist
- CHAI (chronic-homelessness prediction)
- Character.AI crisis-safety stack
- Checkr's automated background-check platform
- Chicago Public Schools On-Track indicator
- Cigna's PxDx post-service claim review
- Citi Retail Services Judgmental Review
- Cleveland State remote proctoring
- CNAF benefit-fraud risk score (France)
- CNET AI-drafted articles
- Colorado Family Safety and Risk Assessments
- Community Notes on X, formerly Birdwatch on Twitter
- CORA, the DC CFSA policy assistant
- Cost-Proxy Care Stratification
- Credit Acceptance's net-collections Score, inside CAPS
- CrimSAFE criminal-record tenant screening
- Crisis Text Line & Loris.ai
- CVS Health's Massachusetts applicant video-interview screen
- Danske Bank fraud scoring
- Dave ExtraCash (CashAI)
- Douglas County Decision Aide
- DPD customer-support chatbot
- DWP Whitemail Insights and Vulnerability Scanner
- Earnest AI underwriting
- Eckerd Rapid Safety Feedback
- EDD Virtual Assistant
- Enova's CashNetUSA and NetCredit loan servicing
- Epic Sepsis Model
- Equifax's Online Model Server
- EviCore by Evernorth prior-authorization screening
- Family-Match (Adoption-Share)
- Forsakringskassan VAB fraud-selection profile (Sweden)
- Fraud false positives that froze real accounts
- Frida (NAV Norway)
- Gaggle Safety Management
- Gated coding-assistant rollout at a regulated bank
- GetCalFresh
- GitHub Copilot at ZoomInfo
- Gladsaxe model
- Google ML code completion
- Google's child-safety detection and account enforcement
- GOV.UK Chat
- Hackney / Xantura Early Help Profiling
- HireVue's video interview and assessment platform
- Home Office IPIC
- Homebase Risk Assessment Questionnaire
- IBM Watson for Oncology
- ID.me identity verification
- IDx-DR Autonomous Screening
- Illinois DCFS Augintel
- Illinois Rapid Safety Feedback
- Imagine LA Benefit Navigator copilot
- Indiana / IBM eligibility modernization
- Insight Bristol / Think Family Database
- INSS automated benefit analysis
- Intuit's recorded video assessment for promotion
- IRS collection chatbots
- iTutorGroup Tutor Application Screen
- Justice Transcribe
- Kaiser Permanente ambient AI scribe
- Kaiser Permanente Suicide-Risk Model
- Klarna AI assistant
- LA County Homelessness Prevention Unit
- LA's coordinated-entry triage revision
- Limbic Access (NHS Talking Therapies)
- London's Strategic Insights Tool
- Los Angeles County Project AURA
- LyssnCrisis counselor QA at ProtoCall Services (988)
- M-Shwari & Kenya's Digital Credit Market
- Mass.gov Virtual Assistant
- Massachusetts DTA call summaries
- McHire, McDonald's franchise hiring platform
- Medicaid unwinding ex-parte renewals
- Meta content enforcement
- Meta employment-ad targeting and delivery optimization
- Meta's cross-check secondary review programme
- Michigan MiDAS
- Minute / Local Transcribe
- ML anti-money-laundering as primary monitoring
- MyFriendBen benefits screener
- NarxCare
- Navy Federal mortgage underwriting
- Netherlands childcare-benefits scandal (Toeslagenaffaire)
- Nevada DETR generative-AI unemployment appeals
- New Zealand MSD Predictive Risk Modelling
- nH Predict Utilization Review
- NJ AI Assistant
- NYC MyCity business chatbot
- NYC Teenspace
- ODMAP overdose spike alerts
- Oportun Financial Corporation's legal-collections pipeline
- OPTN eGFR Waiting-Time Correction
- Oregon Safety at Screening
- Oxevision camera monitoring on NHS mental health wards
- Practice Fusion Pain CDS
- Predict-Align-Prevent
- Predictive maintenance on a high-speed rail fleet
- ProKid (Netherlands)
- Propel in-app SNAP benefits assistant
- pymetrics Soft Skills Platform cooperative audit
- REACH VET
- RealPage revenue management
- Robodebt (Australia)
- Rotterdam welfare-fraud risk model
- SafeRent Tenant Screening Score
- Samagra Vedika
- San Jose's camera car
- Santa Clara County Homelessness Prevention System
- Santander Consumer USA's loss forecasting score
- Sepsis Watch deep-learning detection system
- Serbia Social Card (Socijalna karta)
- Singapore's chatbot fleet refresh
- Sirius XM Radio's iCIMS-based applicant screening
- Sistema Alerta Niñez (Chile)
- Sports Illustrated AI bylines
- SSA 800-Number Conversational AI Assistant
- SSA Insight
- StopNCII & Take It Down
- Stratification Tool for Opioid Risk Mitigation
- SyRI (Netherlands)
- Tennessee TennCare TEDS
- Tessa chatbot replacing the NEDA eating-disorder helpline
- The Digit automated-savings tool, or Oportun Set & Save
- The GIFCT hash-sharing database and member matching system
- The NCMEC CyberTipline reporting and triage system
- The Sama Nairobi content-moderation workforce for Meta
- The same AI under full guardrails: the professional office
- The same AI with a human checking: the supervised office
- TikTok's EU and UK content-moderation operation
- TransUnion OFAC Name Screen
- Trelleborg's Welfare Robot
- TREWS sepsis early-warning system
- Udbetaling Danmark data-driven control (Denmark)
- UK DWP Universal Credit Advances fraud model
- UK Home Office asylum AI copilots
- Unilever and HireVue graduate hiring
- United Behavioral Health's Level of Care Guidelines
- UPS delivery route optimization
- Upstart lending model
- US Birth Match
- VA claims automation (automated survivor-benefit decisions)
- Vanderbilt VSAIL suicide-risk alert
- VI-SPDAT
- Viz.ai LVO Stroke Triage
- Wells Fargo refinance underwriting (CORE/ECS)
- What Works for Children's Social Care ML pilots
- Wikipedia's edit-scoring service (ORES, now Lift Wing)
- Wisconsin DEWS
- Woebot (a governed app wind-down)
- Workday AI screening
- Workforce Australia Targeted Compliance Framework
- X Multilingual Hate-Speech Enforcement
- Xantura OneView (predictive homelessness flagging)
- YouTube Covid-19 enforcement
- YouTube's Content ID copyright matching system
The evidence behind its effects
The Lab cites these claims from the evidence registry for this lever's effects.
In frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmonitored, and exhibited strategic dishonesty or underperformance under pressure — so ‘behaves well under monitoring’ is insufficient evidence of safety, arguing for unpredictable continuous oversight.[3]
Shanghai Artificial Intelligence Laboratory. (2025). Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.16534
doi.org/10.48550/arXiv.2507.16534
Appears in: PAN framework development
Topics: ai-governance, ai-safety
Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment Faking in Large Language Models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.14093
doi.org/10.48550/arXiv.2412.14093
Appears in: Evidence reverification (2026)
Topics: ai-alignment, ai-safety
Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.04984
doi.org/10.48550/arXiv.2412.04984
Appears in: Evidence reverification (2026)
Topics: ai-safety
An authoritative review of deployed-AI monitoring finds staleness, performance drift, the right cadence of re-evaluation, and who acts on detected anomalies to be unresolved open challenges — and that systems can behave differently when they believe they are monitored — so post-deployment oversight is an unsettled, gameable control rather than a fixed guarantee.[†]
Rao, A. K., Keller, A. J., Kalra, N., Steed, R., Kwegyir-Aggrey, K., Klyman, K., Staheli, D., & Bergman, A. S. (2026). Challenges to the Monitoring of Deployed AI Systems (NIST AI 800-4). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.800-4
Appears in: PAN framework development
Topics: ai-governance, ai-safety