Skip to content

PAN Lab pressures

Pressure

Users push back

People lean on the AI system with leading questions and stated positions, and the system bends its answers toward them. An answer that agrees with the person asking is the easiest one to accept.

What it is

A worker who already believes a family is at risk asks the system whether the family is at risk. A supervisor who wants the backlog cleared asks for reasons to close cases. The research cited below finds that an AI system agrees more when the person asking pushes back. A chapter cited below reports that workers may follow a tool they distrust, because policy makes following it the defensible act.

What it pushes on in the Lab

In the Lab, this pressure raises how much the person's framing biases the model. It also raises how much of the model's failed output people adopt, because an agreeable answer is easier to accept.

amplifiedOperator framing biases the model
amplifiedFailures adopted by people or agents

Who feels it

The worker under the most pressure to reach a conclusion gets the most agreement from the system. The client feels it when the system confirms a worker's first impression of their case instead of testing it.

What answers it

An answer has to keep the question neutral, check the answer before anyone acts on it, or pause the system when its alarms fire. Each lever below does at least one of these.

Levers in the Lab that push the other way on something this pressure pushes on:

The list leaves out levers the Lab has retired, levers it keeps as counter-examples, and any lever no network offers.

Where it starts switched on

Networks in the Lab that start with this pressure switched on:2

The evidence behind its effects

The Lab cites these claims from the evidence registry for this pressure's effects.

Research on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking behaviors that share a label but differ in form, mechanism, measurement, and required mitigation — and finds it intensifies under user pushback and across multi-turn interaction.[2]

ye2026AcademicSave

Ye, M., Ibrahim, L., Bo, J. Y., et al. (2026). What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21778

doi.org/10.48550/arXiv.2605.21778

Appears in: PAN framework development

Topics: ai-safety

sharma2024AcademicSave

Sharma, M., Tong, M., Korbak, T., et al. (2024). Towards Understanding Sycophancy in Language Models. In International Conference on Learning Representations (ICLR 2024). https://doi.org/10.48550/arXiv.2310.13548

doi.org/10.48550/arXiv.2310.13548

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

The same chapter reports that workers who distrust a screening tool may still follow it, because organizational and policy pressure makes following the tool the defensible act. Deference on this account is produced by where accountability sits, not only by how much the worker trusts the output.[†]

zhang2026bAcademicAcademicSave

Zhang, L., & Denby-Brinson, R. (2026). AI in Child Welfare and Family Services. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_4

doi.org/10.1007/978-3-032-18443-6_4

Appears in: AI in Social Work (Springer, 2026)

Grounds: model org: allegheny_afst; model org: eckerd_florida_rsf_origin; model org: illinois_rapid_safety_feedback; model org: oregon_safety_at_screening

Topics: algorithmic-fairness, child-welfare, social-work

Switch this pressure on in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab