multimodalfactneutralInstruction tuning for speech language models is substantially more challenging than for text-based LLMs because it requires learning a new modality and speech-specific instructionsComputation and Language27 Jul 2026http://arxiv.org/abs/2607.02214v1