rlhf
fact
neutral
In DPO, the β parameter non-monotonically affects policy deviation at fixed learning rates, with a dead zone at small β, a peak at intermediate values, and decreasing deviation at larger β
at a fixed learning rate the achieved policy deviation is non-monotone in $β$: it vanishes in a dead zone at small $β$, reaches a peak at an intermediate value, and decreases again for larger $β$
Machine Learning29 Aug 2026