Voice of Reason: Reinforcement Learning for Spoken Math
Abstract
Speech language models enable richer spoken interactions between humans and machines than cascaded system, allowing access to paralinguistic information and lower latency. However, their accuracy on mathematical thinking benchmarks have been lagging behind those of text models. Reinforcement learning (RL) with verifiable rewards has been instrumental in extending text models capabilities for solving complex problems, and limiting hallucinations. In this work, we explore applying it to the GLM-4-Voice speech model to bridge the gap between textual and spoken mathematical problem solving. We first adapt the model to the domain using supervised fine-tuning on synthesized audio question answering. We then show that, even without extra reasoning tokens, RL improves the accuracy on GSM8K beyond levels previously achieved for speech models only with supplementary reasoning traces. When combined with existing streaming reasoning techniques, we show further gains to 75.8% free-form accuracy. This establishes a new state-of-the-art for mathematical spoken abilities with speech native models.