rlhf
fact
neutral
Standard DPO loss values are not comparable across different β values, with runs having nearly identical loss curves potentially differing several-fold in KL divergence from the reference model
standard DPO loss values are not comparable across $β$: runs with nearly identical loss curves can differ several-fold in KL divergence from the reference model
Machine Learning29 Aug 2026