HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 1-20 of 81 claims in topic "scaling" of type "fact"

scaling
fact
Neutral
academic

Pre-training with limited data exhibits different scaling behavior than web-scale language modeling, where increasing parameters beyond an optimal point causes overfitting rather than improved performance

"Pre-training under limited data requires a different view of scaling than web-scale language modeling. With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scale; beyond that point, models overfit and generalization worsens."
Machine Learning
8/29/2026
Confidence: 85%Source
2345
scaling
fact
Neutral
academic

Optimal model size depends strongly on both the data budget and the downstream target task

"We study this behavior across 10M-100M word pre-training budgets, two corpora, and multiple downstream evaluations, and find that optimal size depends strongly on both the data budget and the downstream target."
Machine Learning
8/29/2026
Confidence: 90%Source
scaling
fact
Bullish
academic

Recursive Transformers with shared blocks and factorized embeddings outperform standard Transformers on limited data budgets of 10M and 100M words

"We train three recursive models and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners."
Machine Learning
8/29/2026
Confidence: 85%Source
scaling
fact
Bearish
critic

AI industry has current revenues in tens of billions or low hundreds of billions against Capex in the trillions

"we are talking about current revenues in the tens of billions (or low hundreds of billions if you are really optimistic), against Capex in the trillions"
Gary Marcus
8/28/2026
Confidence: 75%Source
scaling
fact
Bullish
journalist

Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made

"Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made. Not gradually, and not evenly — each flip has a patient zero, a paper or product where the synthetic version first became load-bearing at a frontier lab, and from there on, the future is simply here but not yet productionized."
swyx & Alessio
8/28/2026
Confidence: 85%Source
scaling
fact
Bullish
journalist

The reward signal was the first thing to go synthetic in 2022 with InstructGPT and Constitutional AI

"The first thing to go synthetic was, counterintuitively, the judge. InstructGPT established the now-canonical trick: collect human preferences once, train a reward model, and let the policy optimize against the model rather than the humans."
swyx & Alessio
8/28/2026
Confidence: 90%Source
scaling
fact
Bullish
journalist

Advanced skills like finding software vulnerabilities require carrying long causal chains of 20+ inference steps and do not live in total parameter count once a knowledge-holding threshold is reached

"He notes that advanced skills (e.g., finding software vulnerabilities) are not retrieval/memorization problems. They require carrying long causal chains (20+ inference steps) without losing the thread. This ability does not live in total parameter count once a certain knowledge-holding threshold is reached."
swyx & Alessio
8/28/2026
Confidence: 75%Source
scaling
fact
Bullish
journalist

Ornith-1.5 achieves end-to-end self-improvement where the model proposes tasks, generates scaffolds, and produces RL rollouts to create new training experiences

"The headline claim is end-to-end self-improvement: the model proposes tasks, generates scaffolds, and produces RL rollouts to create new training experiences."
swyx & Alessio
8/28/2026
Confidence: 70%Source
scaling
fact
Bearish
academic

More ambiguous next-token distributions are harder for LLMs to learn accurately

"we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately."
Neural and Evolutionary Computing
8/28/2026
Confidence: 90%Source
scaling
fact
Bearish
academic

More ambiguous distributions require more capacity, larger embeddings, more training steps, and amplify sampling noise

"More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise."
Neural and Evolutionary Computing
8/28/2026
Confidence: 85%Source
scaling
fact
Bullish
academic

Scaling post-training alone was sufficient to achieve GLM-5.3's improvements over GLM-5.2 using the same base model

"Scaling post-training is all we did for GLM-5.3"
Nathan Lambert
8/28/2026
Confidence: 85%Source
scaling
fact
Neutral
academic

Larger datasets can make subtle teacher-specific signals easier to detect in trained students, even when examples are off-task

"We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait."
Computation and Language
8/28/2026
Confidence: 85%Source
scaling
fact
Bullish
academic

Larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior

"Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior."
Computation and Language
8/28/2026
Confidence: 90%Source
scaling
fact
Bearish
lab researcher

Base LLM scaling reached a capability plateau between 2022-2024

Francois Chollet
8/9/2026
Confidence: 80%Source
scaling
fact
Bullish
academic

In 2023 and early 2024, Chollet underestimated the long-term importance of LLMs

Francois Chollet
8/8/2026
Confidence: 90%Source
scaling
fact
Bearish
academic

The early 2023 narrative that scaling up base LLMs alone could solve AGI did not pan out

Francois Chollet
8/8/2026
Confidence: 85%Source
scaling
fact
Bearish
academic

Current base LLMs still do not perform well on ARC 1 and can't even reliably do simple math operations

Francois Chollet
8/8/2026
Confidence: 90%Source
scaling
fact
Bearish
critic

Pure scaling of AI models did not work as a strategy for achieving AI progress

Gary Marcus
8/8/2026
Confidence: 90%Source
scaling
fact
Bullish
lab researcher

RL improves reliability more than coverage

Lewis Tunstall
8/8/2026
Confidence: 75%Source
scaling
fact
Neutral
lab researcher

At low total compute budgets, pretraining dominates because weaker policies benefit less from RL

Lewis Tunstall
8/8/2026
Confidence: 80%Source
Page 1 of 5
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.