Skip to content

PAN Lab pressures

Pressure

Autonomy expands

Leadership lets the AI system do more without a person's sign-off. The system writes more directly into the record, and the people around it check less of what it does.

What it is

Autonomy grows one permission at a time. A pilot that drafted notes for review starts filing them. A tool that suggested a next step starts taking it. Each change looks small to the person who approves it.

What it pushes on in the Lab

In the Lab, this pressure raises the failures the system writes directly into records. It also raises operator deference drift, the slide from checking the system's output to accepting it by default.

increasedFailures written directly into records
increasedOperator deference drift

Who feels it

The workers who used to sign off no longer read the entry before the system files it. The client whose record receives an unchecked entry carries whatever error it holds.

What answers it

An answer has to put a person back between the system and the record, cut how much the system writes there, or keep people's habit of checking alive. Each lever below does at least one of these.

Levers in the Lab that push the other way on something this pressure pushes on:

The list leaves out levers the Lab has retired, levers it keeps as counter-examples, and any lever no network offers.

Where it starts switched on

The evidence behind its effects

The Lab cites these claims from the evidence registry for this pressure's effects.

A preprint benchmark reports an in-context misalignment dose-response: in the most susceptible frontier model, up to ~24% misaligned behavior at 16 examples rising to ~58% at 256 examples (rates at 16 examples span roughly 1–24% across models), with the majority of misaligned responses rationalized.[†]

afonin2026AcademicSave

Afonin, N., Andriianov, N., Hovhannisyan, V., Bageshpura, N., Liu, K., Zhu, K., Dev, S., Panda, A., Rogov, O., Tutubalina, E., Panchenko, A., & Seleznyov, M. (2026). Emergent misalignment via in-context learning: Narrow in-context examples can produce broadly misaligned LLMs [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11288

doi.org/10.48550/arXiv.2510.11288

Appears in: Paramerge authored research

Topics: ai-alignment, complexity-science

Switch this pressure on in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab