用隐式运动编码实现高清实时人脸说话头生成
IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation
- 通过隐式运动编码将人脸映射为带外观感知的压缩潜空间
- 支持512x512分辨率下45帧/秒实时生成,性能优于现有扩散与显式模型
- 引入运动统计量捕捉细微口部动作,可调节运动强度与画质平衡
我们提出一种从单张图像和音频输入生成高分辨率说话头的新方法。以往使用显式人脸模型(如3DMM和面部关键点)的方法因缺乏外观感知的运动表征,难以生成高质量视频;而视频扩散模型虽画质高,但处理速度慢,限制实际应用。本文提出的隐式人脸运动扩散模型(IF-MDM)采用隐式运动编码,将人脸嵌入带外观感知的压缩面部潜变量中,提升生成效果。尽管隐式运动缺乏显式模型的空间解耦性,影响微小口部动作对齐,我们引入运动统计量以捕捉精细运动信息。此外,模型在推理时具备运动可控性,可优化运动强度与视觉质量间的权衡。IF-MDM支持512x512分辨率视频实时生成,最高达45帧/秒。大量实验表明其性能显著优于现有扩散模型与显式人脸模型。代码将公开,视频结果见 https://bit.ly/ifmdm_supplementary。
原文摘要 · Abstract (English)
We introduce a novel approach for high-resolution talking head generation from a single image and audio input. Prior methods using explicit face models, like 3D morphable models (3DMM) and facial landmarks, often fall short in generating high-fidelity videos due to their lack of appearance-aware motion representation. While generative approaches such as video diffusion models achieve high video quality, their slow processing speeds limit practical application. Our proposed model, Implicit Face Motion Diffusion Model (IF-MDM), employs implicit motion to encode human faces into appearance-aware compressed facial latents, enhancing video generation. Although implicit motion lacks the spatial disentanglement of explicit models, which complicates alignment with subtle lip movements, we introduce motion statistics to help capture fine-grained motion information. Additionally, our model provides motion controllability to optimize the trade-off between motion intensity and visual quality during inference. IF-MDM supports real-time generation of 512x512 resolution videos at up to 45 frames per second (fps). Extensive evaluations demonstrate its superior performance over existing diffusion and explicit face models. The code will be released publicly, available alongside supplementary materials. The video results can be found on https://bit.ly/ifmdm_supplementary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。