rlhf
fact
bullish
The proposed method consistently outperforms vanilla RLVR across multiple mathematical benchmarks, with performance gains becoming more pronounced as k scales up
Experiments across multiple mathematical benchmarks show that our method consistently outperforms vanilla RLVR. Notably, the performance gain becomes increasingly pronounced as $k$ scales up, demonstrating a substantial expansion of reasoning coverage.
Computation and Language30 Aug 2026