multimodal
fact
bullish
A training-free inference framework can accelerate both the vision encoder and LLM via a unified Condense-and-Extract paradigm that addresses visual token pruning limitations
we propose PACE (Pixel-Adaptive Condense and Extract), a training-free inference framework that accelerates both the vision encoder and the Large Language Model (LLM) via a unified Condense-and-Extract paradigm.
Computer Vision30 Aug 2026