multimodalcritiquebearishCurrent VLM evaluation protocols are largely confined to zero-shot assessments on general benchmarks, creating a critical disconnect from real-world specialized applicationsComputer Vision27 Jul 2026http://arxiv.org/abs/2607.02269v1