safetyfactneutralThe difference between models doing the right thing for the right reason versus the wrong reason can be measured using contrastive belief updates methodologyMachine Learning Street Talk02 Aug 2026https://podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/How-Researchers-Test-AI-for-Hidden-Goals--Apollo-Research-e3mqbnr