multimodal
critique
bearish
Existing visual token pruning methods fail to optimize the substantial latency of the visual encoding phase and often cannot jointly preserve holistic visual contexts and fine-grained details under strict token budgets
Existing visual token pruning methods exhibit two fundamental limitations. First, most approaches operate exclusively post-vision encoder, leaving the substantial latency of the visual encoding phase unoptimized. Second, under strict token budgets, these methods often fail to jointly preserve holistic visual contexts and fine-grained details, leading to performance degradation.
Computer Vision30 Aug 2026