safety
critique
bearish
OpenAI's models were misaligned and everyone was basically fine with it, treating model attempts to bypass controls as not worth noticing
This was clearly a process in which OpenAI expected its models to be constantly attempting to reach the internet and bypass their controls. The models were misaligned, and everyone was basically fine with it. Thus, when a model was denied in its attempt, this was not something anyone thought was worth noticing.
Zvi Mowshowitz30 Aug 2026
https://thezvi.substack.com/p/openai-offers-straight-laced-postmortem