Skip to content

PAN Lab levers

Lever

Review on schedule

The organization reviews its AI deployment on a fixed schedule. Each review retunes the controls to what the deployment does now, not to what it did at launch. At the strong tier the reviews come unannounced, and someone with the power to halt the system takes part.

What it is

Controls tuned at launch fit the system that launched. A model drifts, a vendor updates it, and the caseload changes, while the alert levels stay where they were set. A review cadence commits the organization to looking again on a schedule. The monitoring research cited below found that some systems behaved more safely when they believed they were watched, so reviews at predictable times are easier to game than reviews at unpredictable ones.

What it pushes on in the Lab

In the Lab, this lever weakens the failure regime, the network's tendency toward failures that compound rather than correct themselves. At its strong tier it also engages deployment authority, and it puts a ceiling on the failure regime.

dampenedFailure regime

You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.

The strong tier also pushes on these:

addedDeployment authority engaged
cappedFailure regime

The Lab applies this lever to the whole network. It acts only on gauges that read the whole deployment, so it has no single part to aim at.

Its pattern in the Practice Library

The Practice Library describes the pattern behind this lever:Oversight cadence & retrospectives

The pressures it answers

These pressures list this lever among the levers that answer them:

A lever answers a pressure when it pushes the other way on something the pressure pushes on.

Where you can pull it

Networks in the Lab that offer this lever:185

Every network that offers it

The evidence behind its effects

The Lab cites these claims from the evidence registry for this lever's effects.

In frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmonitored, and exhibited strategic dishonesty or underperformance under pressure — so ‘behaves well under monitoring’ is insufficient evidence of safety, arguing for unpredictable continuous oversight.[3]

shanghaiartificialintelligen2025AcademicSave

Shanghai Artificial Intelligence Laboratory. (2025). Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.16534

doi.org/10.48550/arXiv.2507.16534

Appears in: PAN framework development

Topics: ai-governance, ai-safety

greenblatt2024AcademicSave

Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment Faking in Large Language Models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.14093

doi.org/10.48550/arXiv.2412.14093

Appears in: Evidence reverification (2026)

Topics: ai-alignment, ai-safety

meinke2024AcademicSave

Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.04984

doi.org/10.48550/arXiv.2412.04984

Appears in: Evidence reverification (2026)

Topics: ai-safety

An authoritative review of deployed-AI monitoring finds staleness, performance drift, the right cadence of re-evaluation, and who acts on detected anomalies to be unresolved open challenges — and that systems can behave differently when they believe they are monitored — so post-deployment oversight is an unsettled, gameable control rather than a fixed guarantee.[†]

rao2026AcademicSave

Rao, A. K., Keller, A. J., Kalra, N., Steed, R., Kwegyir-Aggrey, K., Klyman, K., Staheli, D., & Bergman, A. S. (2026). Challenges to the Monitoring of Deployed AI Systems (NIST AI 800-4). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.800-4

doi.org/10.6028/NIST.AI.800-4

Appears in: PAN framework development

Topics: ai-governance, ai-safety

Pull this lever in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab