用隐空间偏差生成语音驱动的手势,提升视频真实感。
Audio-driven Gesture Generation via Deviation Feature in the Latent Space
- 基于隐空间的弱监督学习,捕捉像素级运动偏差。
- 采用扩散模型融合运动特征,生成更细腻的手势与口型。
- 无需精细标注,适合影视动画等真实场景应用。
手势对增强语伴交流至关重要,能提供视觉强调并补充语言互动。以往研究多关注点级动作或完全监督的数据驱动方法,本文聚焦语伴手势,倡导弱监督学习与像素级运动偏差建模。提出一种弱监督框架,通过学习隐空间中的表示偏差,实现语伴手势视频生成。该方法利用扩散模型整合隐含运动特征,实现更精确、更细腻的动作表达。借助隐空间中的弱监督偏差,有效生成手部动作和嘴部运动,显著提升视频真实感。实验表明,该方法在视频质量上超越现有最先进技术。
原文摘要 · Abstract (English)
Gestures are essential for enhancing co-speech communication, offering visual emphasis and complementing verbal interactions. While prior work has concentrated on point-level motion or fully supervised data-driven methods, we focus on co-speech gestures, advocating for weakly supervised learning and pixel-level motion deviations. We introduce a weakly supervised framework that learns latent representation deviations, tailored for co-speech gesture video generation. Our approach employs a diffusion model to integrate latent motion features, enabling more precise and nuanced gesture representation. By leveraging weakly supervised deviations in latent space, we effectively generate hand gestures and mouth movements, crucial for realistic video production. Experiments show our method significantly improves video quality, surpassing current state-of-the-art techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。