rlhffactbearishUniform-weight RL methods like GRPO are suboptimal for HPC tasks due to extreme heterogeneity, with tasks differing by 58x in answer length and spanning three distinct reward distributionsMachine Learning02 Aug 2026http://arxiv.org/abs/2607.28301v1