arXiv:2601.18451cs.CVcs.AI2026-01被引 3

用机器人控制思路生成连贯自然的说话手势,让身体动作与语音更匹配。

3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control

  • 将手势生成转为连续轨迹控制,避免动作碎片化。
  • 在BEAT2数据集上生成手势更自然、语音对齐度更高。
  • 融合语音音素信息,提升表情与动作的语义一致性。

现有方法在生成包含全身动作和面部表情的完整说话手势时,常因分部件或逐帧回归导致动作语义不连贯、空间不稳定。本文提出3DGesPolicy,一种基于机器人运动控制思想的动作驱动框架,将整体手势生成重构为扩散策略下的连续轨迹控制问题。通过统一建模帧间变化为整体动作,有效学习帧间手势运动模式,确保空间与语义上的连贯性,符合真实运动流形。为进一步提升表达对齐,设计了语音-手势-音素(GAP)融合模块,深度整合多模态信号,实现语音语义、身体动作与面部表情的结构化精细对齐。在BEAT2数据集上的大量定量与定性实验表明,3DGesPolicy在生成自然、富有表现力且高度语音对齐的完整手势方面优于现有最先进方法。

原文摘要 · Abstract (English)

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or frame-level regression methods, We introduce 3DGesPolicy, a novel action-based framework that reformulates holistic gesture generation as a continuous trajectory control problem through diffusion policy from robotics. By modeling frame-to-frame variations as unified holistic actions, our method effectively learns inter-frame holistic gesture motion patterns and ensures both spatially and semantically coherent movement trajectories that adhere to realistic motion manifolds. To further bridge the gap in expressive alignment, we propose a Gesture-Audio-Phoneme (GAP) fusion module that can deeply integrate and refine multi-modal signals, ensuring structured and fine-grained alignment between speech semantics, body motion, and facial expressions. Extensive quantitative and qualitative experiments on the BEAT2 dataset demonstrate the effectiveness of our 3DGesPolicy across other state-of-the-art methods in generating natural, expressive, and highly speech-aligned holistic gestures.

手势生成多模态扩散模型语音对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。