multimodalfactbearishCurrent vision-language models fail to infer shared concepts from sets of example images and apply them to new inputsComputer Vision27 Jul 2026http://arxiv.org/abs/2607.02402v1