rlhffactbullishVanilla OPSD's brittleness stems from it being the β=1 case of a broader policy-optimization family, and introducing β as a controllable parameter yields better regularizationMachine Learning02 Aug 2026http://arxiv.org/abs/2607.28582v1