arXiv:2507.05092cs.CV2025-07被引 1

用3D模型和扩散Transformer生成更自然的说话人脸动画。

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

  • 结合3DMM与扩散Transformer,分层去噪提升面部连贯性。
  • 通过3D系数约束实现唇音同步,减少身份漂移,眨眼更自然。
  • 适合虚拟助手、游戏、影视中需要逼真口型同步的场景。

音频驱动的说话人脸生成对虚拟助手、游戏和电影至关重要,自然的唇动是关键。现有方法基于GAN或UNet-based扩散模型,存在三大问题:(i) 时间约束弱导致帧间抖动,造成不一致;(ii) 缺乏足够的3D信息提取,引发身份漂移;(iii) 眨眼行为不自然,缺乏真实眨眼动态建模。为此,我们提出MoDiT,将3D可变形模型(3DMM)与基于扩散的Transformer结合。贡献包括:(i) 提出分层去噪策略,改进时间注意力与偏置自/交叉注意力机制,增强唇音同步并逐步提升全脸一致性,有效缓解时间抖动;(ii) 融合3DMM系数提供显式空间约束,确保3D感知光流预测,利用Wav2Lip结果提升唇音同步,改善身份一致性;(iii) 优化眨眼策略,实现更平滑自然的眼部运动。

原文摘要 · Abstract (English)

Audio-driven talking head generation is critical for applications such as virtual assistants, video games, and films, where natural lip movements are essential. Despite progress in this field, challenges remain in producing both consistent and realistic facial animations. Existing methods, often based on GANs or UNet-based diffusion models, face three major limitations: (i) temporal jittering caused by weak temporal constraints, resulting in frame inconsistencies; (ii) identity drift due to insufficient 3D information extraction, leading to poor preservation of facial identity; and (iii) unnatural blinking behavior due to inadequate modeling of realistic blink dynamics. To address these issues, we propose MoDiT, a novel framework that combines the 3D Morphable Model (3DMM) with a Diffusion-based Transformer. Our contributions include: (i) A hierarchical denoising strategy with revised temporal attention and biased self/cross-attention mechanisms, enabling the model to refine lip synchronization and progressively enhance full-face coherence, effectively mitigating temporal jittering. (ii) The integration of 3DMM coefficients to provide explicit spatial constraints, ensuring accurate 3D-informed optical flow prediction and improved lip synchronization using Wav2Lip results, thereby preserving identity consistency. (iii) A refined blinking strategy to model natural eye movements, with smoother and more realistic blinking behaviors.

说话人脸3DMM扩散模型唇音同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。