让语音模型边说边思考,实时生成逻辑清晰的口语回答。
Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
- 将思考与说话交织在词元层面,打破先想后说的延迟瓶颈。
- 在语音问答任务上提升19.1%算术推理能力,零解码延迟。
- 适合需要实时交互的语音助手、教育对话系统等场景。
推理对有效沟通和决策至关重要。尽管大语言模型和多模态模型已证明显式推理能显著提升理解与泛化能力,但语音语言模型中的推理仍处于初级阶段。早期方法尝试将文本模型的‘先思考后说话’范式迁移至语音,但其顺序结构引入显著延迟,影响实时交互效率。为此,我们提出Mini-Omni-Reasoner框架,通过创新的‘边说边思考’机制,在词元层面上将无声推理词元与有声响应词元交替生成。该设计允许在生成语音的同时嵌入结构化内部推理,利用模型高频词元处理能力实现连续输出。虽为交错结构,但强制局部语义对齐,确保每个响应词元均受前序推理内容指导。为支持该框架,我们构建了大规模数据集Spoken-Math-Problems-3M,确保语音词元始终紧跟相关推理内容,促进语音耦合推理的精准高效学习。基于分层思维者-说话者架构,Mini-Omni-Reasoner实现了流畅且逻辑严谨的口语输出,在Spoken-MQA基准上,算术推理提升19.1%,上下文理解提升6.4%,输出更短,解码延迟为零。
原文摘要 · Abstract (English)
Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly improves understanding and generalization, reasoning in LSMs remains in a nascent stage. Early efforts attempt to transfer the "Thinking-before-Speaking" paradigm from textual models to speech. However, this sequential formulation introduces notable latency, as spoken responses are delayed until reasoning is fully completed, impairing real-time interaction and communication efficiency. To address this, we propose Mini-Omni-Reasoner, a framework that enables reasoning within speech via a novel "Thinking-in-Speaking" formulation. Rather than completing reasoning before producing any verbal output, Mini-Omni-Reasoner interleaves silent reasoning tokens with spoken response tokens at the token level. This design allows continuous speech generation while embedding structured internal reasoning, leveraging the model's high-frequency token processing capability. Although interleaved, local semantic alignment is enforced to ensure that each response token is informed by its preceding reasoning. To support this framework, we introduce Spoken-Math-Problems-3M, a large-scale dataset tailored for interleaved reasoning and response. The dataset ensures that verbal tokens consistently follow relevant reasoning content, enabling accurate and efficient learning of speech-coupled reasoning. Built on a hierarchical Thinker-Talker architecture, Mini-Omni-Reasoner delivers fluent yet logically grounded spoken responses, maintaining both naturalness and precision. On the Spoken-MQA benchmark, it achieves a +19.1% gain in arithmetic reasoning and +6.4% in contextual understanding, with shorter outputs and zero decoding latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。