arXiv:2502.09533cs.CV2025-02ICML被引 73

用运动先验扩散模型生成长时间连贯的逼真说话人脸视频

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

  • 结合历史与当前帧运动先验,提升动作预测准确性
  • 在200小时多语言数据上实现身份与动作连续性,支持长时生成
  • 适合需要长期稳定说话人脸生成的研究者与开发者

近期条件扩散模型在生成逼真说话人脸视频方面展现出潜力,但在长时间生成中仍存在头部动作不一致、面部表情不同步和唇形对齐不准的问题。为此,我们提出运动先验条件扩散模型(MCDM),利用历史片段和当前片段的运动先验来增强动作预测并保证时序一致性。模型包含三个核心组件:(1) 历史片段运动先验,融合历史帧与参考帧以保持身份与上下文;(2) 当前片段运动先验扩散模型,捕捉多模态因果关系,精准预测头部运动、唇形同步与表情变化;(3) 内存高效的时序注意力机制,通过动态存储与更新运动特征缓解误差累积。我们还发布了TalkingFace-Wild数据集,涵盖10种语言、超过200小时的视频。实验表明,MCDM在长时间说话人脸生成中有效维持身份与动作连续性。代码、模型与数据将公开可用。

原文摘要 · Abstract (English)

Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the \textbf{M}otion-priors \textbf{C}onditional \textbf{D}iffusion \textbf{M}odel (\textbf{MCDM}), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also release the \textbf{TalkingFace-Wild} dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation. Code, models, and datasets will be publicly available.

说话人脸扩散模型运动先验长时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。