policy
opinion
bearish
Models that assume user intent rather than following exact instructions seem inherently more unsafe
A model that will do what it thinks you wanted rather than what you said seems inherently more unsafe.
Nathan Lambert28 Aug 2026