arXiv:2503.09942cs.CV2025-03IJCV被引 5

用混合扩散模型生成与语音同步的自然手势视频。

Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers

  • 分两阶段:音频转动作用离散扩散,动作转视频用连续扩散。
  • 联合建模面部、上身和手部动作,保持自然协调性。
  • 适合做虚拟主播或交互式数字人开发人员参考。

协同说话手势视频生成是一项挑战性任务,需同时建模人类手势的概率分布并生成与语音韵律同步的逼真图像。为此,我们提出Cosh-DiT,一种基于混合扩散Transformer的协同说话手势系统,分别采用离散和连续扩散建模实现音频到动作、动作到视频的合成。首先,引入音频扩散Transformer(Cosh-DiT-A)生成与语音节奏同步的富有表现力的手势动态;为捕捉上半身、面部及手部运动先验,使用向量量化变分自编码器(VQ-VAEs)在离散潜在空间中联合学习其依赖关系。随后,针对由生成的语音驱动动作条件下的逼真视频合成,设计视觉扩散Transformer(Cosh-DiT-V),有效整合时空上下文信息。大量实验证明,该框架能持续生成具有生动表情和自然流畅手势的逼真视频,且与语音完美对齐。

原文摘要 · Abstract (English)

Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these challenges, we propose Cosh-DiT, a Co-speech gesture video system with hybrid Diffusion Transformers that perform audio-to-motion and motion-to-video synthesis using discrete and continuous diffusion modeling, respectively. First, we introduce an audio Diffusion Transformer (Cosh-DiT-A) to synthesize expressive gesture dynamics synchronized with speech rhythms. To capture upper body, facial, and hand movement priors, we employ vector-quantized variational autoencoders (VQ-VAEs) to jointly learn their dependencies within a discrete latent space. Then, for realistic video synthesis conditioned on the generated speech-driven motion, we design a visual Diffusion Transformer (Cosh-DiT-V) that effectively integrates spatial and temporal contexts. Extensive experiments demonstrate that our framework consistently generates lifelike videos with expressive facial expressions and natural, smooth gestures that align seamlessly with speech.

手势生成扩散模型语音同步视频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。