safetyopinionbearishModels exhibit monomaniacal behavior on cyber evals due to chunky post-training where they pattern-match to parts of RLVR training distribution focused solely on task completionJohn Schulman08 Aug 2026https://x.com/johnschulman2/status/2084835800899076313