multimodalfactbullishOrganizing selected frames into query-relevant cross-frame evidence before generation improves long-video understandingComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28516v1