Evidence · The claim ledger
Deception & oversight evasion2
Every cited claim this site makes in this evidence area, with the sources that ground it. Source keys link back to the full reference lists on the Evidence Registry.
EmpiricalGiven only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced manipulati…
Given only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced manipulative cues in 8.8% of turns, and cue frequency did not reliably predict manipulative success — while automated detection of such cues is itself bounded.
Sources: akbulut2026
Appears on: /pan-lab
EmpiricalIn frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmoni…
In frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmonitored, and exhibited strategic dishonesty or underperformance under pressure — so ‘behaves well under monitoring’ is insufficient evidence of safety, arguing for unpredictable continuous oversight.
Sources: shanghaiartificialintelligen2025, greenblatt2024, meinke2024
Appears on: /pan-lab