scaling
fact
bearish
More ambiguous next-token distributions are harder for LLMs to learn accurately
we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately.
Neural and Evolutionary Computing28 Aug 2026