arXiv:2507.12804cs.CV2025-07

用音频驱动人脸动画,实现高精度实时生成

ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion

  • 通过声纹转面部关键点,引导噪声扩散提升同步性
  • 在MEAD和CREMA-D数据集上优于现有方法,近实时生成
  • 适合虚拟助手、教育场景等对细节要求高的应用

音频驱动的人脸动画生成需精准同步面部动作与音频信号。本文提出ATL-Diff,通过三个核心组件解决同步难题并降低噪声与计算开销:1)关键点生成模块将音频转换为面部关键点;2)基于关键点的噪声引导机制,按关键点分布分配噪声,解耦音频信息;3)3D身份扩散网络保留身份特征。在MEAD和CREMA-D数据集上的实验表明,ATL-Diff在所有指标上均超越当前最优方法。该方法实现近实时处理,生成高质量动画,有效保留面部细微特征。研究成果可广泛应用于虚拟助手、在线教育、医疗沟通及数字平台。代码已开源:https://github.com/sonvth/ATL-Diff

原文摘要 · Abstract (English)

Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and computational costs. Our framework features three key components: a Landmark Generation Module converting audio to facial landmarks, a Landmarks-Guide Noise approach that decouples audio by distributing noise according to landmarks, and a 3D Identity Diffusion network preserving identity characteristics. Experiments on MEAD and CREMA-D datasets demonstrate that ATL-Diff outperforms state-of-the-art methods across all metrics. Our approach achieves near real-time processing with high-quality animations, computational efficiency, and exceptional preservation of facial nuances. This advancement offers promising applications for virtual assistants, education, medical communication, and digital platforms. The source code is available at: \href{https://github.com/sonvth/ATL-Diff}{https://github.com/sonvth/ATL-Diff}

音频驱动人脸生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。