HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claimsrlhf
rlhf
critique
bearish

On-policy distillation effectiveness depends on teacher consistency, where the OPD supervision model should have generated the SFT demonstrations, but this condition is frequently violated in practice

Computation and Language02 Aug 2026

http://arxiv.org/abs/2607.28449v1