benchmarksfactbearishKnowledge retrieval and reasoning is the primary bottleneck in Knowledge-Intensive Visual Question Answering, not visual grounding or object identificationComputer Vision28 Jul 2026http://arxiv.org/abs/2607.21155v1