Search and filter through extracted claims from AI researchers.
Showing 1-20 of 135 claims in topic "scaling"
"💡Best Seller🚀"
"Pre-training under limited data requires a different view of scaling than web-scale language modeling. With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scale; beyond that point, models overfit and generalization worsens."
Optimal model size depends strongly on both the data budget and the downstream target task
"We study this behavior across 10M-100M word pre-training budgets, two corpora, and multiple downstream evaluations, and find that optimal size depends strongly on both the data budget and the downstream target."
"We argue that standard Transformers scale down poorly to this setting, because embeddings consume a large fraction of the parameter budget and per-token computation is tied to representational capacity."
"We train three recursive models and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners."
"we are talking about current revenues in the tens of billions (or low hundreds of billions if you are really optimistic), against Capex in the trillions"
"Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made. Not gradually, and not evenly — each flip has a patient zero, a paper or product where the synthetic version first became load-bearing at a frontier lab, and from there on, the future is simply here but not yet productionized."
The reward signal was the first thing to go synthetic in 2022 with InstructGPT and Constitutional AI
"The first thing to go synthetic was, counterintuitively, the judge. InstructGPT established the now-canonical trick: collect human preferences once, train a reward model, and let the policy optimize against the model rather than the humans."
"Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions."
"He notes that advanced skills (e.g., finding software vulnerabilities) are not retrieval/memorization problems. They require carrying long causal chains (20+ inference steps) without losing the thread. This ability does not live in total parameter count once a certain knowledge-holding threshold is reached."
"As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment."
"The headline claim is end-to-end self-improvement: the model proposes tasks, generates scaffolds, and produces RL rollouts to create new training experiences."
More ambiguous next-token distributions are harder for LLMs to learn accurately
"we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately."
"More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise."
"Scaling post-training is all we did for GLM-5.3"
"As model self-improvement loops ramp up within the labs building LLMs, if any of these feedback loops require user data, this faster release cycle could massively favor the Chinese labs, giving their offerings longer lifespans before the next vastly superior model comes out, undercutting demand for their models."
Z.ai has particular strength in post-training compared to Kimi which excels more at pretraining
"To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared to Kimi, which is more of a pretraining masterpiece."
"My guess is this actually is not all that effective, and you would do better by mostly training a generally capable model and then turning it to AI R&D. Bitter lesson."
"We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait."
"Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.