rlhffactneutralRemoving GRPO's standard deviation normalization transforms catastrophic reward performance from 0% to baseline parityMachine Learning28 Jul 2026http://arxiv.org/abs/2607.21273v1