interpretabilityfactneutralArchitecture prior, not expressivity, determines which algorithm transformers learn: weight tying causes models to select serial frontiers instead of parallel scansMachine Learning (Statistics)28 Jul 2026http://arxiv.org/abs/2607.20594v1