safetycritiquebearishCurrent safety alignment efforts inadvertently alter models' representations of mindedness in other entities alongside human beliefs and valuesComputation and Language02 Aug 2026http://arxiv.org/abs/2607.28607v1