multimodalfactbullishRefCaptioner achieves best overall performance among open-source models for multi-reference image-grounded video captioningComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28509v1