Search and filter through extracted claims from AI researchers.
Showing 1-20 of 233 claims in topic "reasoning"
"Difficult concepts such as softmax, temperature, and top-p sampling are clarified with code-linked explanations and diagrams, while visual workflows make pipelines and scoring methods easier to follow."
The book offers a guided, project-driven learning experience rather than a broad survey.
"Reading the book feels like following a guided technical build rather than a loose survey of AI topics."
Refinement methods can sometimes degrade answers, a common failure mode in reasoning models.
"The book also discusses common failure modes, including cases where refinement can make answers worse."
"Readers see how self-consistency, self-refinement, Best-of-N, and training-based methods actually work, including their cost and latency trade-offs."
The book implements core reasoning methods from scratch rather than using black-box library calls.
"The book is especially useful because it implements the core methods from scratch rather than treating them as black-box library calls."
Knowledge graphs and LLMs can be used together to build AI systems using connected data.
""Knowledge Graphs and LLMs in Action: Build AI systems using connected data""
"At @OpenAI we aim to devote time to finding and announcing math results from internal models only when they would meaningfully change people’s understanding of the pace of AI progress."
"Our main focus is shipping great models so everyone can use them to make discoveries of their own."
"Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost"
"CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples"
LLM failure modes exhibit structured patterns across model scales within the same family
"LLM failure modes exhibit structured patterns across model scales within the same family"
"Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs)"
"rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct"
"Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token"
TTPO shows strong cross-task generalization capabilities
"shows strong cross-task generalization"
TTPO yields +25.2% to +36.4% improvement without thinking
"yields +25.2% to +36.4% without thinking"
TTPO raises Qwen3-1.7B model performance from 38.0% to 45.2% during test-time training
"raises Qwen3-1.7B from 38.0% to 45.2% in TTT"
"Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks"
"Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models"
"these methods typically rely on repeated generation or external verification"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.