interpretabilityfactneutralAdapter memorization capacity depends more on where parameters sit than on parameter count, with MLP locations holding nearly twice as much as attentionMachine Learning28 Jul 2026http://arxiv.org/abs/2607.21351v1