multimodalfactneutralVisual tokens in Omni-LLMs are highly redundant and account for the vast majority of input costComputer Vision28 Jul 2026http://arxiv.org/abs/2607.21179v1