让对话机器人同时说与动,实现自然互动
DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

- 双塔Transformer架构,语音与动作流式同步生成
- 在4000小时数据上训练,跨说话人依赖建模更准确
- 适合虚拟助手、社交机器人等实时交互场景
我们提出DyaPlex,一种面向二人互动的流式全双工语音-动作模型。为捕捉人类交流中持续且互惠的特性,该模型具备同时感知并生成语音与身体动作的能力。核心方法基于一个基础全双工语音模型的强先验,并引入新型动作路径,实现多模态完全同步。具体地,采用双塔Transformer结构,在保留冻结语音模型零样本对话推理能力的同时,构建深度耦合的流式动作路径。通过统一的二人组标记交错机制及时间对齐的语音-动作RoPE引导跨注意力,模型有效将自回归动作与丰富的语音隐状态对齐。在4000小时的Seamless Interaction数据集上训练后,模型成功捕捉跨说话人依赖关系,在单人与双人互动基准测试中均达到新最佳性能。
原文摘要 · Abstract (English)
We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this full-duplex capability empowers the agent to simultaneously perceive and generate both speech and physical motion in a streaming fashion. At its core, our method leverages the strong priors of a foundational full-duplex speech model and integrates a novel motion pathway, thereby achieving fully synchronized multi-modal interaction. Specifically, we design a dual-tower Transformer architecture that preserves the zero-shot conversational reasoning of a frozen base speech model while constructing a deeply coupled, streaming motion pathway. By introducing a unified dyadic token interleaving mechanism and guiding cross-attention via a time-aligned speech-motion RoPE, our model effectively aligns autoregressive motions with rich latent speech features. Trained on the 4,000-hour Seamless Interaction dataset, our model effectively captures cross-speaker dependencies and establishes new state-of-the-art performance across both monadic and dyadic human interaction benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。