safetycritiquebearishPrompt-level safety evaluation hides important failures in models, as they often fail to remain safe across matched intent variantsComputation and Language27 Jul 2026http://arxiv.org/abs/2607.02047v1