Search and filter through extracted claims from AI researchers.
Showing 101-120 of 2685 claims of type "fact"
LLM failure modes exhibit structured patterns across model scales within the same family
"LLM failure modes exhibit structured patterns across model scales within the same family"
"Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs)"
"Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities"
Persistent knowledge accumulation in the wiki is critical for effective skill evolution
"persistent knowledge accumulation in the wiki is critical for effective skill evolution"
"evolved skills transfer effectively across models and model families, and skills evolved by other models can outperform self-evolved skills"
"larger models generally benefit more from evolved skills, while smaller models with skills can outperform substantially larger models without them"
There are now 38 identifiable patterns of clichéd language produced by LLMs
"My LLM cliché highlighter is up to 38 patterns now"
"Across diverse benchmarks and models, WikiSkill consistently outperforms state-of-the-art skill-evolution methods"
"rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct"
TTPO shows strong cross-task generalization capabilities
"shows strong cross-task generalization"
TTPO yields +25.2% to +36.4% improvement without thinking
"yields +25.2% to +36.4% without thinking"
TTPO raises Qwen3-1.7B model performance from 38.0% to 45.2% during test-time training
"raises Qwen3-1.7B from 38.0% to 45.2% in TTT"
"Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks"
"Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models"
"the first stage performs trajectory-level screening based on process quality, result quality, and data representativeness, selecting a high-quality and representative subset of successful trajectories."
"training on the 10% trajectory subset selected by SWE-Prime outperforms training on the full resolved dataset, yielding relative performance gains of up to 12.2% and 24.2%, respectively."
"task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective, redundant, or risky steps. Directly using such trajectories for SFT can introduce noisy supervision and encourage models to imitate undesirable problem-solving behaviors."
"our in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory"
"LLMs' performance varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed"
"experiments reveal that mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.