Skip to content

PAN Lab levers

Lever

Check with a second model

The organization has a genuinely different second model check the first model's output. Where the two disagree, a person looks closer. A copy of the same model would share the first model's blind spots, so the check depends on a real difference between the two.

What it is

Two copies of one model make the same mistakes at the same time. In the Michigan MiDAS case cited below, one uniform rule set produced tens of thousands of correlated wrongful fraud determinations, one flaw repeated at the scale of the caseload. A second model from another vendor or another design disagrees where the first is wrong in a way the second is not. The research cited below also finds that a model rarely corrects its own answer once it has committed to it. The check therefore has to come from outside the model.

What it pushes on in the Lab

In the Lab, this lever strengthens the pathway by which models cross-check each other's output. It also weakens the pathway by which AI models or agents relay failures to one another. Its side effect raises operator deference drift, because two systems that agree look like proof.

amplifiedModels cross-check each other's outputs
dampenedFailures relayed between AI models or agents
increasedOperator deference driftside effect

You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.

In the modes that offer aiming, you can aim this lever at particular parts and pathways of a network. Otherwise it applies to the whole network.

Its pattern in the Practice Library

The Practice Library describes the pattern behind this lever:Cross-model verification

The pressures it answers

These pressures list this lever among the levers that answer them:

A lever answers a pressure when it pushes the other way on something the pressure pushes on.

Where you can pull it

Networks in the Lab that offer this lever:112

Every network that offers it

The evidence behind its effects

The Lab cites these claims from the evidence registry for this lever's effects.

A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.[2]

Language models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consistent with it — recognizing 67-87% of those fabrications as false when re-asked in a clean, uncontaminated context but not correcting them in place — so one error deterministically spawns supporting errors, a self-sustaining failure the model's own downstream output feeds.[†]

zhang2024AcademicSave

Zhang, M., Press, O., Merrill, W., Liu, A., & Smith, N. A. (2024). How Language Model Hallucinations Can Snowball. In International Conference on Machine Learning (ICML 2024), PMLR 235:59670-59684. https://doi.org/10.48550/arXiv.2305.13534

doi.org/10.48550/arXiv.2305.13534

Appears in: PAN framework development

Topics: ai-safety

Pull this lever in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab