multimodalfactbullishSpeechCombine can create instruction-following speech language models without instruction tuning, using only continuous pre-training on 30k hours of speech dataComputation and Language27 Jul 2026http://arxiv.org/abs/2607.02214v1