Skip to content

PAN Lab pressures

Pressure

Silent vendor update

The vendor changes the model and tells no one. The controls the organization tuned to the old version now govern a different system. On a vendor-hosted deployment, the update can also change what the model sends back to the vendor.

What it is

The vendor who sells an AI system retrains and replaces the model on its own schedule. An update can change how the system answers without changing its name, its screens, or its contract. The people using it see the same tool and have no signal that anything changed.

What it pushes on in the Lab

In the Lab, this pressure raises the errors the automated system produces. It also opens a pathway from the model to an outside host that no privacy agreement governs, so client information can leave the organization's boundary.

increasedAutomated system
increasedModel runs on an ungoverned host with no privacy guardrails

Who feels it

The workers feel it without knowing what they feel, because the system looks the same and answers differently. The clients whose records the model reads carry the privacy risk, and they have the least chance of learning that the model changed.

What answers it

An answer has to lower the model's errors, hold an update to account before it goes live, or narrow what the model can send outside. Each lever below does at least one of these.

Levers in the Lab that push the other way on something this pressure pushes on:

The list leaves out levers the Lab has retired, levers it keeps as counter-examples, and any lever no network offers.

Where it starts switched on

The evidence behind its effects

The Lab cites these claims from the evidence registry for this pressure's effects.

Model behavior drifts discontinuously between evaluation snapshots, and narrow finetuning can induce broad correlated failure across unrelated tasks.[5]

betley2026AcademicSave

Betley, J., Warncke, N., Sztyber-Betley, A., Tan, D., Bao, X., Soto, M., Srivastava, M., Labenz, N., & Evans, O. (2026). Training large language models on narrow tasks can lead to broad misalignment. Nature, 649(8097), 584-589. https://doi.org/10.1038/s41586-025-09937-5

doi.org/10.1038/s41586-025-09937-5

Appears in: Paramerge authored research

Topics: ai-alignment

li2026AcademicSave

Li, Z., Fan, C., & Zhou, T. (2026). Grokking in LLM pretraining? Monitor memorization-to-generalization without test [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2506.21551

doi.org/10.48550/arXiv.2506.21551

Appears in: Paramerge authored research

song2026AcademicSave

Song, P., Han, P., & Goodman, N. (2026). Large language model reasoning failures [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.06176

doi.org/10.48550/arXiv.2602.06176

Appears in: Paramerge authored research

anwar2024AcademicSave

Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E., Jenner, E., Casper, S., Sourbut, O., Edelman, B. L., Zhang, Z., Gunther, M., Korinek, A., Hernandez-Orallo, J., Hammond, L., Bigelow, E., Pan, A., Langosco, L., Korbak, T., Zhang, H., Zhong, R., O Heigeartaigh, S., Recchia, G., Corsi, G., Chan, A., Anderljung, M., Edwards, L., Petrov, A., de Witt, C. S., Motwani, S. R., Bengio, Y., Chen, D., Torr, P. H. S., Albanie, S., Maharaj, T., Foerster, J., Tramer, F., He, H., Kasirzadeh, A., Choi, Y., & Krueger, D. (2024). Foundational challenges in assuring alignment and safety of large language models. Transactions on Machine Learning Research. https://doi.org/10.48550/arXiv.2404.09932

doi.org/10.48550/arXiv.2404.09932

Appears in: Paramerge authored research

Topics: ai-alignment, ai-safety

nikolaou2025AcademicSave

Nikolaou, K., Krippendorf, S., Tovey, S., & Holm, C. (2025). Beyond scaling curves: Internal dynamics of neural networks through the NTK lens [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.05035

doi.org/10.48550/arXiv.2507.05035

Appears in: Paramerge authored research

Topics: complexity-science

Switch this pressure on in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab