arXiv:2605.20356cs.CLcs.AI2026-05

研究双工语音对话模型如何同步与提前预判对方说话,模拟真实对话。

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

论文配图:Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
图 1 · 摘自论文原文
  • 用Moshi模型在控制条件下模拟双工对话,观察内部表征同步性。
  • 无噪声时表征同步最强,随信道噪声增加而下降,且能提前预测换话时机。
  • 揭示模型内部状态可编码换话预判信号,适合研究对话智能的团队参考。

全双工语音对话模型(SDMs)可同时听和说,使交互更接近人类对话,而非传统的轮流模式。受人类交流中神经耦合的启发,我们研究此类模型在交互过程中如何协调其内部表征。在受控条件下,通过两个预训练的Moshi模型模拟全双工对话,操控信道噪声和解码偏差。使用中心化核对齐(CKA)在时间滞后上测量同步性,并利用因果LSTM模型从说话方和听者视角探测延迟的内部激活中蕴含的预期换话线索。结果表明,在无噪声条件下存在强表征同步,峰值出现在零滞后附近,且随着噪声增加而减弱;同时发现内部状态编码了支持提前预测换话的信息。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue models (SDMs) can listen and speak simultaneously, enabling interaction dynamics closer to human conversation than turn-based systems. Inspired by neural coupling in human communication, we study how such models coordinate their internal representations during interaction. We simulate full-duplex dialogues between two instances of the pretrained \textit{Moshi} model under controlled conditions, manipulating channel noise and decoding bias. Synchronization is measured using Centered Kernel Alignment (CKA) across temporal lags, while anticipatory turn-taking cues are probed from delayed internal activations using causal LSTM models, from both speaker and listener perspectives. We find strong representational synchronization under no noise conditions, peaking near zero lag and degrading with noise, and we show that internal states encode anticipatory information that supports turn-taking prediction ahead of time.

对话系统语音模型同步机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。