rlhffactbearishDense per-step supervision with group-normalized RL can destroy policy performance rather than improve itMachine Learning28 Jul 2026http://arxiv.org/abs/2607.21273v1