scaling
critique
bearish
Standard Transformers scale down poorly to limited data settings because embeddings consume too large a fraction of parameters and per-token computation is coupled with representational capacity
We argue that standard Transformers scale down poorly to this setting, because embeddings consume a large fraction of the parameter budget and per-token computation is tied to representational capacity.
Machine Learning29 Aug 2026