safety
fact
bearish
The trained act/don't-act gate is context-fragile and only holds when response format allows reasoning; rigid formats leave models confident and wrong on questions they otherwise answer correctly
It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly.
Computation and Language30 Aug 2026