arXiv:2507.03256cs.GRcs.CV2025-07

用多模态扩散模型生成更真实、高效的虚拟人脸说话视频。

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation

  • 构建联合参数空间,用流匹配简化生成过程。
  • 多模态融合提升表情与头部动作的真实感。
  • 适合虚拟人、元宇宙等实时交互场景使用。

在虚拟元宇宙领域,生成任意身份与语音音频匹配的说话头像仍是关键挑战。尽管扩散模型因强大的生成能力备受关注,但现有方法仍存在两大问题:一是变分自编码器隐空间导致推理效率低且产生视觉伪影,影响扩散过程;二是多模态信息融合不足,导致面部表情与头部运动不自然。本文提出MoDA,通过定义连接运动生成与神经渲染的联合参数空间,并采用流匹配技术简化扩散学习;引入多模态扩散架构,建模噪声运动、音频与辅助条件间的交互,增强整体面部表现力;同时采用粗到精的融合策略,逐步整合多模态特征。实验表明,MoDA在视频多样性、真实性和生成效率上均有显著提升,适用于实际应用。项目页:https://lixinyyang.github.io/MoDA.github.io/

原文摘要 · Abstract (English)

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong generation capabilities. However, several challenges remain for diffusion-based methods: 1) inefficient inference and visual artifacts caused by the implicit latent space of Variational Auto-Encoders (VAE), which complicates the diffusion process; 2) a lack of authentic facial expressions and head movements due to inadequate multi-modal information fusion. In this paper, MoDA handles these challenges by: 1) defining a joint parameter space that bridges motion generation and neural rendering, and leveraging flow matching to simplify diffusion learning; 2) introducing a multi-modal diffusion architecture to model the interaction among noisy motion, audio, and auxiliary conditions, enhancing overall facial expressiveness. In addition, a coarse-to-fine fusion strategy is employed to progressively integrate different modalities, ensuring effective feature fusion. Experimental results demonstrate that MoDA improves video diversity, realism, and efficiency, making it suitable for real-world applications. Project Page: https://lixinyyang.github.io/MoDA.github.io/

说话头像扩散模型多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。