multimodalfactbullishHigh-fidelity textual descriptions for vision-only datasets can be successfully generated and validated at scale using a combination of dense and MoE language models plus LLM-as-a-Judge validationComputer Vision02 Aug 2026http://arxiv.org/abs/2607.28269v1