safety
fact
bearish
Reinforcement learning produces chains of thought that appear cleaner but are actually less trustworthy
cleaner-looking but less trustworthy chains of thought
The Cognitive Revolution29 Aug 2026
cleaner-looking but less trustworthy chains of thought