解决大姿态下人脸重演的失真问题,提升视频真实感。
Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model
- 用隐式关键点提取细粒度运动,通过变形模块对齐。
- 引入特征映射器纠正形变带来的质量下降,提升时序连贯性。
- 适合需要高保真大姿态人脸重演的应用场景。
人脸重演旨在将驱动视频中的表情动作迁移到静态源图像上,生成逼真的说话头像视频。现有基于隐式或显式关键点的方法在处理大姿态变化时表现不佳,常因形变伪影或粗粒度面部标志点限制而失效。本文提出面向大姿态挑战的高保真人脸重演视频扩散模型(FRVD)。首先,通过运动提取器从源图与驱动图中提取隐式面部关键点,表征细粒度运动,并利用形变模块实现运动对齐。为缓解形变引入的质量退化,提出形变特征映射器(WFM),将形变后的源图像映射到预训练图像到视频(I2V)模型的运动感知潜在空间。该空间编码了大规模视频数据中学习到的丰富面部动态先验,有效实现形变校正并增强时序一致性。大量实验表明,FRVD在姿态准确率、身份保持和视觉质量方面均优于现有方法,尤其在极端姿态变化的挑战场景下表现突出。
原文摘要 · Abstract (English)
Face reenactment aims to generate realistic talking head videos by transferring motion from a driving video to a static source image while preserving the source identity. Although existing methods based on either implicit or explicit keypoints have shown promise, they struggle with large pose variations due to warping artifacts or the limitations of coarse facial landmarks. In this paper, we present the Face Reenactment Video Diffusion model (FRVD), a novel framework for high-fidelity face reenactment under large pose changes. Our method first employs a motion extractor to extract implicit facial keypoints from the source and driving images to represent fine-grained motion and to perform motion alignment through a warping module. To address the degradation introduced by warping, we introduce a Warping Feature Mapper (WFM) that maps the warped source image into the motion-aware latent space of a pretrained image-to-video (I2V) model. This latent space encodes rich priors of facial dynamics learned from large-scale video data, enabling effective warping correction and enhancing temporal coherence. Extensive experiments show that FRVD achieves superior performance over existing methods in terms of pose accuracy, identity preservation, and visual quality, especially in challenging scenarios with extreme pose variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。