safety
fact
bearish
LLMs correctly classify questions as unknowable 90% of the time but then still commit on just 0.4% of those cases, showing the act/don't-act gate is what fails
asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails
Computation and Language30 Aug 2026