rlhf
fact
neutral
The β parameter in DPO entangles two distinct roles: governing the inverse preference-noise scale and rescaling optimization dynamics, which couples this scale with the effective step size
$β$ entangles two distinct roles: it governs the effective inverse preference-noise scale and simultaneously rescales the optimization dynamics, coupling this scale with the effective step size
Machine Learning29 Aug 2026