rlhf
fact
bullish
Forcing the target model to generate answers based on partial reasoning trajectories from smaller, weaker language models effectively disrupts over-confidence and encourages exploration of distinct reasoning paths
Instead of relying solely on internal exploration, we force the target model to generate answers based on partial reasoning trajectories generated by a smaller, weaker language models. These unfamiliar prefixes effectively disrupt over-confidence and encourage the exploration of distinct reasoning paths.
Computation and Language30 Aug 2026