multimodalcritiquebearishVision-language models may generate incomplete, erroneous, or misleading scene descriptions in automotive in-car scene understanding applicationsComputer Vision27 Jul 2026http://arxiv.org/abs/2607.02300v1