arXiv:2412.04448cs.CV2024-12被引 43

MEMO通过记忆与情绪感知,让口型、表情更自然连贯。

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

  • 用记忆模块存长期上下文,提升身份一致性和动作平滑度
  • 结合音频情绪检测,使表情更符合语音情感,对齐效果更好
  • 适合做高质量数字人视频生成,尤其注重表情真实性的场景

近期视频扩散模型的发展为逼真音频驱动的说话视频生成带来了新可能。然而,实现无缝的音画唇形同步、保持长期身份一致性以及生成自然且与音频匹配的表情仍是重大挑战。为此,我们提出内存引导的情绪感知扩散模型(MEMO),一种端到端的音频驱动肖像动画方法,用于生成身份一致且富有表现力的说话视频。该方法包含两个核心模块:(1) 内存引导的时间模块,通过线性注意力机制利用存储的长时上下文信息,增强长期身份一致性和运动平滑性;(2) 情绪感知音频模块,将传统交叉注意力替换为多模态注意力以强化音视频交互,并通过音频情绪检测,使用情绪自适应层归一化来优化面部表情。大量定量与定性实验表明,MEMO在多种图像和音频类型下生成的视频更逼真,在整体质量、音唇同步、身份一致性和表情-情绪对齐方面均优于现有最先进方法。

原文摘要 · Abstract (English)

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing natural, audio-aligned expressions in generated talking videos remain significant challenges. To address these challenges, we propose Memory-guided EMOtion-aware diffusion (MEMO), an end-to-end audio-driven portrait animation approach to generate identity-consistent and expressive talking videos. Our approach is built around two key modules: (1) a memory-guided temporal module, which enhances long-term identity consistency and motion smoothness by developing memory states to store information from a longer past context to guide temporal modeling via linear attention; and (2) an emotion-aware audio module, which replaces traditional cross attention with multi-modal attention to enhance audio-video interaction, while detecting emotions from audio to refine facial expressions via emotion adaptive layer norm. Extensive quantitative and qualitative results demonstrate that MEMO generates more realistic talking videos across diverse image and audio types, outperforming state-of-the-art methods in overall quality, audio-lip synchronization, identity consistency, and expression-emotion alignment.

视频生成扩散模型情绪感知身份一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。