用隐式运动迁移实现高效高保真语音驱动人脸生成
IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
- 用交叉注意力替代光流,统一建模运动与身份
- 40帧/秒(视频驱动)和42帧/秒(音频驱动)实时生成
- 适合需要身份保持的虚拟人、短视频生成场景
说话人脸生成旨在从单张图像合成逼真的说话肖像,但现有方法多依赖显式光流和局部变形,难以建模复杂全局运动并导致身份漂移。本文提出IMTalker,通过隐式运动迁移实现高效高保真说话人脸生成。核心思想是用跨注意力机制替代传统光流变形,在统一潜在空间中隐式建模运动差异与身份对齐,实现鲁棒的全局运动渲染。为在跨身份重演中保持说话者身份,引入身份自适应模块,将运动潜在变量投影至个性化空间,确保运动与身份清晰解耦。此外,设计轻量级光流匹配运动生成器,从音频、姿态和注视线索生成生动可控的隐式运动向量。大量实验表明,IMTalker在运动准确性、身份保持和音唇同步方面优于现有方法,达到当前最优质量,且效率优异:在RTX 4090 GPU上实现视频驱动40 FPS、音频驱动42 FPS的实时生成。代码与预训练模型将公开以促进应用与后续研究。
原文摘要 · Abstract (English)
Talking face generation aims to synthesize realistic speaking portraits from a single image, yet existing methods often rely on explicit optical flow and local warping, which fail to model complex global motions and cause identity drift. We present IMTalker, a novel framework that achieves efficient and high-fidelity talking face generation through implicit motion transfer. The core idea is to replace traditional flow-based warping with a cross-attention mechanism that implicitly models motion discrepancy and identity alignment within a unified latent space, enabling robust global motion rendering. To further preserve speaker identity during cross-identity reenactment, we introduce an identity-adaptive module that projects motion latents into personalized spaces, ensuring clear disentanglement between motion and identity. In addition, a lightweight flow-matching motion generator produces vivid and controllable implicit motion vectors from audio, pose, and gaze cues. Extensive experiments demonstrate that IMTalker surpasses prior methods in motion accuracy, identity preservation, and audio-lip synchronization, achieving state-of-the-art quality with superior efficiency, operating at 40 FPS for video-driven and 42 FPS for audio-driven generation on an RTX 4090 GPU. We will release our code and pre-trained models to facilitate applications and future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。