Search and filter through extracted claims from AI researchers.
Showing 21-40 of 233 claims in topic "reasoning"
Evolution Strategies have emerged as a memory-efficient post-training paradigm for LLM reasoning
"Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning."
"We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning."
"The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood."
"The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity."
"this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs."
"Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances."
"Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO."
"we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates."
"This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting."
ES requires a smaller population size in a larger LLM
"we study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM."
"These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO."
Chain-of-thought reasoning can be used to improve the reliability of LLM responses
"Large Language Models (LLMs) can be trained to perform chain-of-thoughts reasoning in order to improve the reliability of their responses."
"We introduce a novel fragment-based reasoning framework in which the model first extracts parallel source-target fragments from retrieved similar exemplars, and uses these fragments as intermediate reasoning traces to produce the final translation."
"Our experiments with the Qwen3 model family, over 6 languages, including up to 5 domains per language, demonstrate that fragment-based MT significantly outperforms alternative methods like standard k-shot or basic drafting."
"batch prompting makes large language model inference more efficient by processing multiple instances simultaneously, it suffers from unpredictable downstream task performance"
"cascaded batch prompting, a two-stage approach designed to resolve the unpredictability of conventional batch prompting by disentangling complex reasoning from symbol grounding"
"the proposed method outperforms the standard single prompting baseline while achieving a speedup proportional to batch size, establishing a new state of the art on the Pareto frontier"
"such RL-style post-training ("RL-for-LLMs") can substantially improve chain-of-thought reasoning, long-horizon planning, and self-correction"
State-of-the-art reasoning language model training requires millions of GPU-hours
"state-of-the-art RLM training requires millions of GPU-hours"
RLM training is as much a parallel and distributed systems problem as an algorithmic one
"This makes RLM training as much a parallel and distributed systems problem as an algorithmic one"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.