让说话时的全身动作更自然,支持语音与文本双重控制。
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
- 用文本到动作数据补充语音到动作数据,解决全身动作缺失问题。
- 通过多阶段训练对齐语音、动作与提示的嵌入空间,实现精准控制。
- 适合需要精细全身动作生成的研究者与动画制作人员。
现有协同语音动作生成方法通常仅关注上半身手势,难以基于文本提示实现说话时行走等复杂全身动作的协同控制。主要挑战在于:1)现有语音到动作数据集包含的全身动作类型有限,超出训练分布;2)缺乏标注的用户提示。为此,我们提出 SynTalker,利用现成的文本到动作数据集作为辅助,补全缺失的全身动作与提示信息。核心技术贡献有二:一是多阶段训练流程,在语音-动作与文本-动作数据分布差异大的情况下,实现动作、语音与提示的嵌入空间对齐;二是基于扩散模型的条件推理过程,采用先分离后组合策略,实现局部肢体的细粒度控制。大量实验表明,该方法可基于语音与用户提示精确灵活地生成协同全身动作,超越现有方法的能力。
原文摘要 · Abstract (English)
Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, such as talking while walking. The major challenges lie in 1) the existing speech-to-motion datasets only involve highly limited full-body motions, making a wide range of common human activities out of training distribution; 2) these datasets also lack annotated user prompts. To address these challenges, we propose SynTalker, which utilizes the off-the-shelf text-to-motion dataset as an auxiliary for supplementing the missing full-body motion and prompts. The core technical contributions are two-fold. One is the multi-stage training process which obtains an aligned embedding space of motion, speech, and prompts despite the significant distributional mismatch in motion between speech-to-motion and text-to-motion datasets. Another is the diffusion-based conditional inference process, which utilizes the separate-then-combine strategy to realize fine-grained control of local body parts. Extensive experiments are conducted to verify that our approach supports precise and flexible control of synergistic full-body motion generation based on both speeches and user prompts, which is beyond the ability of existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。