让语音模型边说边思考,像人一样分脑处理。
Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
- 用两个脑分工:一个想清楚,一个流畅说
- 零延迟下数学题准确率达92.8%,对话得分82.5
- 适合需要实时交互的语音助手与教育应用
实时语音语言模型(SLMs)因需串行生成完整思维链而难以应用思维链(CoT)推理。我们首次提出脑启发式「心速说话」(Mind-Paced Speaking, MPS)框架,实现高保真、实时推理。类比人类不同脑区分别负责思考与回应,MPS采用双脑架构:由「构想脑」进行高层推理以引导「表达脑」生成流畅语音。这种分工消除模式切换,保持推理完整性。实验表明,MPS显著优于现有边想边说方法,在预计算完整思维链模型的推理表现上几乎持平,但延迟大幅降低。在零延迟配置下,该方法在数学推理任务Spoken-MQA上取得92.8%准确率,在语音对话任务URO-Bench上获82.5分。MPS是开源系统Step-Audio R1.1的核心方法,有效弥合高质量推理与实时交互之间的差距。
原文摘要 · Abstract (English)
Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entire thought process sequentially. Enabling SLMs to think while speaking, similar to humans, is attracting increasing attention. We present, for the first time, Mind-Paced Speaking (MPS), a brain-inspired framework that enables high-fidelity, real-time reasoning. Similar to how humans utilize distinct brain regions for thinking and responding, we propose a novel dual-brain approach, employing a "Formulation Brain" for high-level reasoning to pace and guide a separate "Articulation Brain" for fluent speech generation. This division of labor eliminates mode-switching, preserving the integrity of the reasoning process. Experiments show that MPS significantly outperforms existing think-while-speaking methods and achieves reasoning performance comparable to models that pre-compute the full CoT before speaking, while drastically reducing latency. Under a zero-latency configuration, the proposed method achieves an accuracy of 92.8% on the mathematical reasoning task Spoken-MQA and attains a score of 82.5 on the speech conversation task URO-Bench. MPS is the methodology underlying our released Step-Audio R1.1 system, effectively bridging the gap between high-quality reasoning and real-time interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。