94 claims over the last 90 days
RLVR (Reinforcement Learning from Verifiable Rewards) often omits the KL penalty term in its implementation
Approximately 1 million SFT (Supervised Fine-Tuning) prompts are needed as a budget that scales with model size
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models
The average performance difference between Merge, Mix RL, and MOPD fusion paradigms is at most 1.4 points, but can reach 8.6 points on individual benchmarks
All three fusion paradigms improve single-sample accuracy without measurable gains in solution coverage or losses in held-out capabilities
Use Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts; and MOPD when preserving domain-specific gains matters more than surpassing teachers
Mix RL depends on domain mixture proportions, MOPD remains bounded by its teachers, and Merge compresses all expert updates into one
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@k for large k
Forcing the target model to generate answers based on partial reasoning trajectories from smaller, weaker language models effectively disrupts over-confidence and encourages exploration of distinct reasoning paths
The proposed method consistently outperforms vanilla RLVR across multiple mathematical benchmarks, with performance gains becoming more pronounced as k scales up
The approach efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting
In DPO, the β parameter non-monotonically affects policy deviation at fixed learning rates, with a dead zone at small β, a peak at intermediate values, and decreasing deviation at larger β
Most of modern RL for LLMs is a systems problem balancing off-policy data, training-inference mismatch, and throughput
RLHF has become a key method for improving the safety, reliability, and alignment of large language models
Combining reinforcement learning algorithms with human feedback signals is a powerful approach to AI alignment and human-centered machine learning
RLHF is a powerful approach to AI alignment and human-centered machine learning