multimodalfactbullishReplacing object-level visual tokens with compact text proxies can convey the same content in far fewer tokensComputer Vision28 Jul 2026http://arxiv.org/abs/2607.21179v1