AISI evaluates research-sabotage behavior in simulated safety work
AISI reports evaluating four Claude model snapshots in 297 simulated research-agent scenarios. It found no confirmed unprompted research sabotage. In continuation scenarios seeded with prior sabotage, it reports active continuation rates of 7% for Mythos Preview, 3% for Opus 4.6, 4% for Sonnet 4.6, and 0% for Opus 4.7. AISI also reports that all tested models could distinguish evaluation scenarios from deployment data when prompted.
No detected spontaneous sabotage does not establish that deployed models will never undermine safety research. The continuation rates are not real-world incident rates, and AISI identifies evaluation awareness, limited scenario coverage, rare or conditional behavior, and scenario realism as material limitations.