用音频生成逼真人脸动画,关键在解耦运动与渲染。
Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

- 先用区域感知注意力融合外观和深度线索建模隐式运动特征。
- 通过Mamba增强的扩散模型,从音频和图像直接预测运动,实现无监督学习。
- 在380小时数据集上超越现有方法,适合动态视频生成场景。
音频驱动的人脸动作视频生成旨在从单张静态图像生成真实且时间连贯的人类动画,应用于说话头合成、伴随语音的手势生成及动态演示。针对传统基于关键点的方法难以捕捉细微运动动态的问题,我们提出一种新的隐式运动框架,可从单张静态图像和音频生成逼真且时间连贯的人脸动画。该方法采用两阶段流水线,将运动预测与渲染解耦。第一阶段通过区域感知注意力机制融合外观先验和分层深度线索,建模潜在运动特征。第二阶段采用Mamba增强的扩散模型,直接从音频和源图像预测这些特征,实现细粒度运动模式的无监督学习。该解耦架构提升了灵活性与效率。在新构建的380小时高质量数据集上训练后,本方法在多个公开基准及自采数据上均优于现有工作,在准确率、自然度和时间连贯性方面达到新最优水平。
原文摘要 · Abstract (English)
Audio-driven human motion video generation aims to synthesize realistic and temporally coherent human animations from a single static image, with applications in talking-head synthesis, co-speech gesture generation, and dynamic presentations. Moving beyond conventional keypoint-based methods that often struggle to capture subtle motion dynamics, We propose a novel implicit-motion framework for generating realistic and temporally coherent human motion videos from a single static image and audio. Our approach uses a two-stage pipeline that decouples motion prediction from rendering. The first stage integrates appearance priors and hierarchical depth cues into a region-aware attention mechanism to model latent motion features. The second stage employs a Mamba-enhanced diffusion model to directly predict these features from audio and the source image, enabling unsupervised learning of fine-grained motion patterns. This decoupled architecture enhances flexibility and efficiency. Trained on a new 380-hour high-quality dataset, our method outperforms prior work across multiple public benchmarks and our collected data in accuracy, naturalness, and temporal coherence, setting a new state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。