arXiv:2512.13131cs.AIcs.CV2025-12被引 1

通过分层隐周期学习,让语音驱动的3D手势更自然协调。

Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning

  • 用周期自编码器分离手势运动相位,保留真实分布与个性差异。
  • 分层级联引导面部、身体、手部动作,提升多部位协同性。
  • 适合做虚拟人动画、语音驱动角色生成的研究者参考。

从语音生成基于3D的身体动作在下游应用中潜力巨大,但仍难以模拟真实人类动作。现有研究多采用端到端生成方案,如GAN、VQ-VAE和扩散模型,但这些方法未能建模不同动作单元(头部、身体、手部)间的内在相关性,导致动作不自然且协调性差。本文提出统一的分层隐周期(HIP)学习方法,通过两个关键设计:一是利用周期自编码器探索手势运动相位流形,既保留真实动作分布,又融合当前潜在状态引入实例级多样性;二是通过级联引导机制建模面部、身体与手部动作的层次关系,实现协同驱动。在3D虚拟人上的实验表明,该方法在定量与定性评估上均优于现有最佳方法。代码与模型将公开共享。

原文摘要 · Abstract (English)

Generating 3D-based body movements from speech shows great potential in extensive downstream applications, while it still suffers challenges in imitating realistic human movements. Predominant research efforts focus on end-to-end generation schemes to generate co-speech gestures, spanning GANs, VQ-VAE, and recent diffusion models. As an ill-posed problem, in this paper, we argue that these prevailing learning schemes fail to model crucial inter- and intra-correlations across different motion units, i.e. head, body, and hands, thus leading to unnatural movements and poor coordination. To delve into these intrinsic correlations, we propose a unified Hierarchical Implicit Periodicity (HIP) learning approach for audio-inspired 3D gesture generation. Different from predominant research, our approach models this multi-modal implicit relationship by two explicit technique insights: i) To disentangle the complicated gesture movements, we first explore the gesture motion phase manifolds with periodic autoencoders to imitate human natures from realistic distributions while incorporating non-period ones from current latent states for instance-level diversities. ii) To model the hierarchical relationship of face motions, body gestures, and hand movements, driving the animation with cascaded guidance during learning. We exhibit our proposed approach on 3D avatars and extensive experiments show our method outperforms the state-of-the-art co-speech gesture generation methods by both quantitative and qualitative evaluations. Code and models will be publicly available.

手势生成分层建模语音驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。