reasoning
fact
bullish
ES can lead to broader reasoning coverage than GRPO, thereby better exploiting the reasoning capabilities of pretrained LLMs
this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs.
Machine Learning30 Aug 2026