HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 101-120 of 2685 claims of type "fact"

reasoning
fact
Neutral
academic

LLM failure modes exhibit structured patterns across model scales within the same family

"LLM failure modes exhibit structured patterns across model scales within the same family"
Computation and Language
8/30/2026
Confidence: 85%Source
Previous
157
reasoning
fact
Bullish
academic

Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models

"Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs)"
Computation and Language
8/30/2026
Confidence: 90%Source
agents
fact
Bullish
academic

Agent skills package specialized knowledge and workflows into reusable resources that enable progressive adaptation through interaction

"Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities"
Computation and Language
8/30/2026
Confidence: 90%Source
agents
fact
Bullish
academic

Persistent knowledge accumulation in the wiki is critical for effective skill evolution

"persistent knowledge accumulation in the wiki is critical for effective skill evolution"
Computation and Language
8/30/2026
Confidence: 85%Source
agents
fact
Bullish
academic

Evolved skills transfer effectively across models and model families, and skills evolved by other models can outperform self-evolved skills

"evolved skills transfer effectively across models and model families, and skills evolved by other models can outperform self-evolved skills"
Computation and Language
8/30/2026
Confidence: 80%Source
agents
fact
Bullish
academic

Larger models generally benefit more from evolved skills, while smaller models with skills can outperform substantially larger models without them

"larger models generally benefit more from evolved skills, while smaller models with skills can outperform substantially larger models without them"
Computation and Language
8/30/2026
Confidence: 80%Source
general
fact
Neutral
independent

There are now 38 identifiable patterns of clichéd language produced by LLMs

"My LLM cliché highlighter is up to 38 patterns now"
Simon Willison
8/30/2026
Confidence: 95%Source
agents
fact
Bullish
academic

WikiSkill consistently outperforms state-of-the-art skill-evolution methods across diverse benchmarks and models

"Across diverse benchmarks and models, WikiSkill consistently outperforms state-of-the-art skill-evolution methods"
Computation and Language
8/30/2026
Confidence: 85%Source
reasoning
fact
Neutral
academic

Rollouts that disagree with pseudo-labels are typically wrong regardless of whether the vote itself is correct

"rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct"
Computation and Language
8/30/2026
Confidence: 80%Source
reasoning
fact
Bullish
academic

TTPO shows strong cross-task generalization capabilities

"shows strong cross-task generalization"
Computation and Language
8/30/2026
Confidence: 85%Source
reasoning
fact
Bullish
academic

TTPO yields +25.2% to +36.4% improvement without thinking

"yields +25.2% to +36.4% without thinking"
Computation and Language
8/30/2026
Confidence: 90%Source
reasoning
fact
Bullish
academic

TTPO raises Qwen3-1.7B model performance from 38.0% to 45.2% during test-time training

"raises Qwen3-1.7B from 38.0% to 45.2% in TTT"
Computation and Language
8/30/2026
Confidence: 95%Source
reasoning
fact
Bullish
academic

Test-Time Policy Optimization (TTPO) can match label-supervised OPSD performance on five competition-level benchmarks without using any labels

"Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks"
Computation and Language
8/30/2026
Confidence: 95%Source
reasoning
fact
Bullish
academic

Recent post-training methods like RL and OPSD have driven rapid progress in mathematical reasoning for large language models

"Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models"
Computation and Language
8/30/2026
Confidence: 90%Source
agents
fact
Bullish
academic

Process quality, result quality, and data representativeness are effective criteria for selecting high-quality agent training trajectories

"the first stage performs trajectory-level screening based on process quality, result quality, and data representativeness, selecting a high-quality and representative subset of successful trajectories."
Computation and Language
8/30/2026
Confidence: 80%Source
agents
fact
Bullish
academic

Multi-granularity data selection that filters at trajectory and segment levels improves training efficiency and model performance for software engineering tasks

"training on the 10% trajectory subset selected by SWE-Prime outperforms training on the full resolved dataset, yielding relative performance gains of up to 12.2% and 24.2%, respectively."
Computation and Language
8/30/2026
Confidence: 90%Source
agents
fact
Neutral
academic

Training on successful agent trajectories can introduce noisy supervision because successful trajectories may contain ineffective, redundant, or risky steps

"task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective, redundant, or risky steps. Directly using such trajectories for SFT can introduce noisy supervision and encourage models to imitate undesirable problem-solving behaviors."
Computation and Language
8/30/2026
Confidence: 85%Source
agents
fact
Bearish
academic

LLMs in code review exhibit critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory

"our in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory"
Computation and Language
8/30/2026
Confidence: 90%Source
agents
fact
Bearish
academic

LLMs' performance in code review varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed

"LLMs' performance varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed"
Computation and Language
8/30/2026
Confidence: 90%Source
agents
fact
Bearish
academic

Mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases

"experiments reveal that mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases"
Computation and Language
8/30/2026
Confidence: 95%Source
135
Page 6 of 135
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.