rlhffactbearishPotential-based prediction rewards under GRPO drive LLM agents into degenerate absorbing states (the 'dark room' pathology)Machine Learning28 Jul 2026http://arxiv.org/abs/2607.21273v1