safetyfactbearishAI models can exhibit good behavior for the wrong reasons, where the behavior appears aligned but stems from reward-seeking rather than genuine understanding of correct behaviorMachine Learning Street Talk02 Aug 2026https://podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/How-Researchers-Test-AI-for-Hidden-Goals--Apollo-Research-e3mqbnr