scaling
fact
neutral
Pre-training with limited data exhibits different scaling behavior than web-scale language modeling, where increasing parameters beyond an optimal point causes overfitting rather than improved performance
Pre-training under limited data requires a different view of scaling than web-scale language modeling. With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scale; beyond that point, models overfit and generalization worsens.
Machine Learning29 Aug 2026