rlhf
fact
neutral
Any fitted ratio with small calibration error performs nearly as well as the best post-processing based on its fitted values, as shown by a calibration-refinement bound
we derive a calibration-refinement bound showing that any fitted ratio with small calibration error performs nearly as well as the best post-processing based on its fitted values.
Machine Learning (Statistics)29 Aug 2026