arXiv:2607.07148eess.AS2026-07

分离说话时机与内容,让对话模型更自然流畅。

Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning

论文配图:Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning
图 1 · 摘自论文原文
  • 用强化学习将何时说话与说什么分开训练
  • 提升即时回应、抢话处理等交互表现
  • 适合追求自然对话体验的语音助手研发

近期全双工语音对话模型在低延迟响应、产生回应信号和处理用户打断方面取得显著进展,但这些交互能力的提升常以牺牲推理和指令遵循能力为代价,暴露出交互动态与智能能力之间的潜在矛盾。本文认为这种权衡并非本质问题:对话动态可作为独立的实时决策策略,从人类对话数据中学习。为此,我们提出DuplexPO,一种强化学习框架,将说话时机与语义内容解耦。该框架保留指令微调助手的语义响应能力,同时在长对话中选取高影响时段优化其时间交互行为。为量化优化动态,我们设计了因子化对话动态奖励(FCDR),实现对发言启动、回应信号、让步和参与度的细粒度时间信用分配,并采用类似GRPO的目标进行策略优化。实验表明,DuplexPO显著提升全双工行为表现,包括及时回应、顺畅换言及打断处理,同时保持强大的推理与指令遵循性能。动态优化指标的提升也反映在更好的用户体验上,说明将对话时机作为独立目标优化,可促进更自然的全双工交互。

原文摘要 · Abstract (English)

Recent full-duplex spoken dialogue models have demonstrated compelling progress toward human-like interaction, enabling agents to respond with low latency, produce backchannels, and handle user barge-ins. Yet these improvements in conversational dynamics often come with weaker reasoning and instruction-following abilities, revealing a potential tension between interactive dynamics and intelligence capability. In this paper, we argue that such an intelligence--dynamics trade-off is not fundamental: conversational dynamics can instead be learned as a separate real-time decision policy from human dialogue data. To this end, we propose DuplexPO, a reinforcement learning (RL) framework that decouples when to speak from what to say. It preserves the semantic response capability of an instruction-tuned assistant, while optimizing its temporal interaction behavior over selected high-impact windows from long human conversations. To quantitatively optimize these dynamics, we formulate the Factorized Conversational Dynamics Reward (FCDR) to enable fine-grained temporal credit assignment for turn initiation, backchanneling, yielding, and regularized participation. The policy is then optimized with a GRPO-style objective. Experiments show that DuplexPO substantially improves full-duplex behaviors, including timely backchannels, smooth turn-taking, and barge-in handling, while maintaining strong reasoning and instruction-following performance. Moreover, improvements in dynamics-oriented metrics are reflected in better user experience, suggesting that optimizing conversational timing as a standalone objective can promote more natural full-duplex interaction.

对话系统强化学习全双工语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。