HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 1-20 of 129 claims in topic "reasoning" of type "fact"

reasoning
fact
Bullish
academic

The book implements core reasoning methods from scratch rather than using black-box library calls.

"The book is especially useful because it implements the core methods from scratch rather than treating them as black-box library calls."
Kirk Borne
9/1/2026
Confidence: 80%Source
2345
reasoning
fact
Bullish
academic

CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods while requiring significantly fewer generations and lower token cost

"Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost"
Computation and Language
8/30/2026
Confidence: 90%Source
reasoning
fact
Bullish
academic

CritICL improves reasoning while maintaining high efficiency by leveraging failure modes from weaker models as guidance through critique-based in-context examples

"CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples"
Computation and Language
8/30/2026
Confidence: 85%Source
reasoning
fact
Neutral
academic

LLM failure modes exhibit structured patterns across model scales within the same family

"LLM failure modes exhibit structured patterns across model scales within the same family"
Computation and Language
8/30/2026
Confidence: 85%Source
reasoning
fact
Bullish
academic

Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models

"Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs)"
Computation and Language
8/30/2026
Confidence: 90%Source
reasoning
fact
Neutral
academic

Rollouts that disagree with pseudo-labels are typically wrong regardless of whether the vote itself is correct

"rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct"
Computation and Language
8/30/2026
Confidence: 80%Source
reasoning
fact
Bullish
academic

TTPO shows strong cross-task generalization capabilities

"shows strong cross-task generalization"
Computation and Language
8/30/2026
Confidence: 85%Source
reasoning
fact
Bullish
academic

TTPO yields +25.2% to +36.4% improvement without thinking

"yields +25.2% to +36.4% without thinking"
Computation and Language
8/30/2026
Confidence: 90%Source
reasoning
fact
Bullish
academic

TTPO raises Qwen3-1.7B model performance from 38.0% to 45.2% during test-time training

"raises Qwen3-1.7B from 38.0% to 45.2% in TTT"
Computation and Language
8/30/2026
Confidence: 95%Source
reasoning
fact
Bullish
academic

Test-Time Policy Optimization (TTPO) can match label-supervised OPSD performance on five competition-level benchmarks without using any labels

"Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks"
Computation and Language
8/30/2026
Confidence: 95%Source
reasoning
fact
Bullish
academic

Recent post-training methods like RL and OPSD have driven rapid progress in mathematical reasoning for large language models

"Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models"
Computation and Language
8/30/2026
Confidence: 90%Source
reasoning
fact
Bullish
academic

Evolution Strategies have emerged as a memory-efficient post-training paradigm for LLM reasoning

"Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning."
Machine Learning
8/30/2026
Confidence: 85%Source
reasoning
fact
Bullish
academic

Synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning has a global finite-sample guarantee with leading last-iterate fluctuation of order Õ(T^(-a/2)/√(1-γ)) for stepsizes α_t=c(t+1)^(-a) with a∈(1/2,1), with no polynomial dependence on the number of quantiles.

"We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning."
Machine Learning (Statistics)
8/30/2026
Confidence: 95%Source
reasoning
fact
Neutral
academic

The proof technique separates two stability mechanisms: a global comparison argument based on order monotonicity and W_∞ contraction brings an arbitrarily initialized iterate into a local neighborhood, where linearization and martingale analysis apply.

"The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood."
Machine Learning (Statistics)
8/30/2026
Confidence: 90%Source
reasoning
fact
Bullish
academic

The result sharply distinguishes between local stochastic fluctuation and global sample complexity in distributional RL, with the deterministic transient and burn-in depending on the smallest Bellman-target density (order m^(-1) in worst case).

"The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity."
Machine Learning (Statistics)
8/30/2026
Confidence: 90%Source
reasoning
fact
Bullish
academic

ES can lead to broader reasoning coverage than GRPO, thereby better exploiting the reasoning capabilities of pretrained LLMs

"this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs."
Machine Learning
8/30/2026
Confidence: 80%Source
reasoning
fact
Bullish
academic

Verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances

"Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances."
Machine Learning
8/30/2026
Confidence: 85%Source
reasoning
fact
Bullish
academic

ES improves Pass@1 while attaining higher Pass@K than GRPO, unlike GRPO which exhibits entropy collapse

"Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO."
Machine Learning
8/30/2026
Confidence: 80%Source
reasoning
fact
Neutral
academic

The task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates despite substantial whole-model parameter drift

"we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates."
Machine Learning
8/30/2026
Confidence: 75%Source
reasoning
fact
Bullish
academic

Large parameter movement in ES need not imply widespread functional change and does not necessarily lead to catastrophic forgetting

"This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting."
Machine Learning
8/30/2026
Confidence: 75%Source
6
7
Page 1 of 7
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.