scaling
fact
bullish
Recursive Transformers with shared blocks and factorized embeddings outperform standard Transformers on limited data budgets of 10M and 100M words
We train three recursive models and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners.
Machine Learning29 Aug 2026