arXiv:2606.11167cs.CLeess.AS2026-06中稿 · EMNLP

用强化学习提升全双工语音模型的对话互动性,解决沉默和抢话问题。

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

论文配图:Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
图 1 · 摘自论文原文
  • 通过四类交互行为设计专属奖励函数,分轴优化对话表现。
  • 在真实对话中显著减少沉默、改善抢话与回应及时性。
  • 适用于需自然交互的语音助手、虚拟伴侣等实时对话场景。

全双工语音对话模型可同时听与说,是实现自然对话的有前途架构。然而,当前模型仅通过令牌级似然最大化进行监督学习,未直接优化交互层面行为,导致沉默过多、发言时机不当等问题。尽管已有研究引入强化学习(RL)改善互动性,但现有方法仅针对有限的交互行为设计奖励。本文提出一种后训练对齐方法,通过强化学习全面提升全双工语音对话模型的互动性。针对四个经典交互维度——停顿处理、发言切换、附和回应和用户打断——从人类对话语料中提取短音频片段,并使用特定奖励函数优化模型。额外引入基于大语言模型的响应质量奖励,防止语义退化。我们在两个开源模型Moshi和PersonaPlex上应用该方法,在离线预录音频评估与实时多轮对话评估中均实现一致的互动性提升。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-level likelihood maximization, which does not directly optimize interaction-level behaviors, causing interactivity issues such as excessive silence and ill-timed turn-taking. Recent work has applied reinforcement learning (RL) to improve interactivity, but existing methods address only a limited set of interactive behaviors in their rewards. In this work, we propose a post-training alignment method that comprehensively improves the interactivity of full-duplex spoken dialogue models through RL. We address the four canonical axes of interactivity: pause handling, turn-taking, backchanneling, and user interruption. For each axis, we extract short audio segments from human conversation corpora and optimize the model with axis-specific reward functions. An extra LLM-based reward for response quality prevents semantic degradation. We apply our method to two open-source models, Moshi and PersonaPlex, demonstrating consistent improvements in interactivity on both offline evaluation with pre-recorded audio and real-time multi-turn dialogue evaluation.

语音对话强化学习全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。