实时生成两人对话时的互动动作,支持在线响应。
It Takes Two: Real-time Co-Speech Two-person's Interaction Generation via Reactive Auto-regressive Diffusion Model
- 基于扩散模型的自回归系统,融合双方状态与语音输入
- 在多任务测试中优于现有方法,支持动态交互生成
- 首个可在线生成双人互动动作的系统,适合虚拟助手等场景
对话场景在现实世界中十分常见,但现有双人共说话动作合成方法往往无法有效捕捉一方语音与姿态对另一方反应的影响。此外,多数方法依赖离线序列到序列框架,难以用于在线应用。本文提出一种语音驱动的自回归系统,用于合成两人对话时的动态全身动作。核心是基于扩散的全身动作生成模型,条件包括双方历史状态、语音音频及任务导向的动作轨迹输入,实现灵活的空间控制。为增强模型学习多样化互动的能力,我们扩充了现有双人对话动作数据集,加入更多动态交互动作。通过多项实验验证,该系统在单人与双人共说话动作生成以及交互动作生成任务中均表现更优。据我们所知,这是首个能够从语音实时生成双人互动全身动作的系统。
原文摘要 · Abstract (English)
Conversational scenarios are very common in real-world settings, yet existing co-speech motion synthesis approaches often fall short in these contexts, where one person's audio and gestures will influence the other's responses. Additionally, most existing methods rely on offline sequence-to-sequence frameworks, which are unsuitable for online applications. In this work, we introduce an audio-driven, auto-regressive system designed to synthesize dynamic movements for two characters during a conversation. At the core of our approach is a diffusion-based full-body motion synthesis model, which is conditioned on the past states of both characters, speech audio, and a task-oriented motion trajectory input, allowing for flexible spatial control. To enhance the model's ability to learn diverse interactions, we have enriched existing two-person conversational motion datasets with more dynamic and interactive motions. We evaluate our system through multiple experiments to show it outperforms across a variety of tasks, including single and two-person co-speech motion generation, as well as interactive motion generation. To the best of our knowledge, this is the first system capable of generating interactive full-body motions for two characters from speech in an online manner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。