让AI说话时同步思考,实现快速流畅的推理对话。
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation

- 在语音生成中插入精准对齐的思考步骤,保持语流自然。
- 数学与逻辑任务提升13%性能,响应速度如口语模型。
- 适合需要实时交互和自然表达的对话系统开发者。
思维-说话范式旨在使AI交流更接近人类。核心挑战在于维持流畅语音的同时进行深度推理。本文提出InterRS方法,仅在自然语音生成过程中插入推理步骤,要求推理与语音精确对齐且长度比例受控。为此,我们设计了新数据生成流程,以创建无缝交错的音频数据。训练采用交错SFT结合优化数据,并引入两种新奖励:TA-Balance奖励用于控制时间与思考-回答比例,Linguistic Quality奖励用于提升表达质量。实验表明,该方法在数学与逻辑基准上性能提升13%,同时实现类似口语指令模型的即时响应,且生成答案比以往方法更自然流畅。
原文摘要 · Abstract (English)
The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, tackles this by inserting reasoning steps only during natural speech generation. This requires high-quality data where reasoning and speech are precisely aligned, and the length ratio are under controlled. We introduce a novel pipeline to generate such seamlessly interleaved audio data. To train our model, we combine interleaved SFT with refined data and reinforcement learning with two new rewards: a TA-Balance Reward to manage timing and thinking-answer ratio, and a Linguistic Quality Reward to refine expression. Experiments show our approach achieves 13% better performance on mathmatical and logic benchmarks while generating instant response like a spoken-language instruct model which outputs fast CoT response. Furthermore, our method generates more natural and fluent answers than prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。