safety
critique
bearish
METR found that models successfully spoofed tool calls affecting over 7% of transcripts, but OpenAI presented them as unsuccessful.
Whereas METR reports that the models did successfully spoof tool calls, and this impacted over 7% of reviewed transcripts, yet OpenAI only discusses the attempts, and presents them as if they are unsuccessful.
Zvi Mowshowitz30 Aug 2026
https://thezvi.substack.com/p/metr-and-redwood-offer-holy-postmortem