Search and filter through extracted claims from AI researchers.
Showing 1-20 of 81 claims in topic "scaling" of type "fact"
"Pre-training under limited data requires a different view of scaling than web-scale language modeling. With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scale; beyond that point, models overfit and generalization worsens."
Optimal model size depends strongly on both the data budget and the downstream target task
"We study this behavior across 10M-100M word pre-training budgets, two corpora, and multiple downstream evaluations, and find that optimal size depends strongly on both the data budget and the downstream target."
"We train three recursive models and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners."
"we are talking about current revenues in the tens of billions (or low hundreds of billions if you are really optimistic), against Capex in the trillions"
"Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made. Not gradually, and not evenly — each flip has a patient zero, a paper or product where the synthetic version first became load-bearing at a frontier lab, and from there on, the future is simply here but not yet productionized."
The reward signal was the first thing to go synthetic in 2022 with InstructGPT and Constitutional AI
"The first thing to go synthetic was, counterintuitively, the judge. InstructGPT established the now-canonical trick: collect human preferences once, train a reward model, and let the policy optimize against the model rather than the humans."
"He notes that advanced skills (e.g., finding software vulnerabilities) are not retrieval/memorization problems. They require carrying long causal chains (20+ inference steps) without losing the thread. This ability does not live in total parameter count once a certain knowledge-holding threshold is reached."
"The headline claim is end-to-end self-improvement: the model proposes tasks, generates scaffolds, and produces RL rollouts to create new training experiences."
More ambiguous next-token distributions are harder for LLMs to learn accurately
"we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately."
"More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise."
"Scaling post-training is all we did for GLM-5.3"
"We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait."
"Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior."
Base LLM scaling reached a capability plateau between 2022-2024
In 2023 and early 2024, Chollet underestimated the long-term importance of LLMs
The early 2023 narrative that scaling up base LLMs alone could solve AGI did not pan out
Pure scaling of AI models did not work as a strategy for achieving AI progress
At low total compute budgets, pretraining dominates because weaker policies benefit less from RL
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.