multimodalfactbullishA training framework and architecture that learns to infer visual concepts from image sets generates more accurate and diverse outputs and generalizes to unseen conceptsComputer Vision27 Jul 2026http://arxiv.org/abs/2607.02402v1