multimodalfactbearishSpeech input amplifies partial failures in multimodal models compared to text inputspeech amplifying partial failuresComputation and Language30 Aug 2026http://arxiv.org/abs/2608.27135v1