Search and filter through extracted claims from AI researchers.
Showing 1-20 of 129 claims in topic "reasoning" of type "fact"
The book implements core reasoning methods from scratch rather than using black-box library calls.
"The book is especially useful because it implements the core methods from scratch rather than treating them as black-box library calls."
"Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost"
"CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples"
LLM failure modes exhibit structured patterns across model scales within the same family
"LLM failure modes exhibit structured patterns across model scales within the same family"
"Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs)"
"rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct"
TTPO shows strong cross-task generalization capabilities
"shows strong cross-task generalization"
TTPO yields +25.2% to +36.4% improvement without thinking
"yields +25.2% to +36.4% without thinking"
TTPO raises Qwen3-1.7B model performance from 38.0% to 45.2% during test-time training
"raises Qwen3-1.7B from 38.0% to 45.2% in TTT"
"Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks"
"Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models"
Evolution Strategies have emerged as a memory-efficient post-training paradigm for LLM reasoning
"Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning."
"We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning."
"The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood."
"The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity."
"this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs."
"Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances."
"Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO."
"we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates."
"This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.