首个基于扩散模型的视频人脸替换框架,提升一致性与细节质量
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
- 融合静态图与视频数据的混合训练,增强时序连贯性
- 在多个指标上优于现有方法,且推理步数更少
- 适合需要高保真人脸替换的影视合成与虚拟角色应用
视频人脸替换在诸多应用中日益流行,但现有方法多集中于静态图像,难以应对视频场景下的时序一致性与复杂情况。本文提出首个专为视频人脸替换设计的扩散模型框架。该框架采用图像-视频混合训练策略,结合大量静态图像数据与时间序列视频,克服纯视频训练的局限性。引入专门设计的扩散模型与VidFaceVAE,有效处理两类数据,更好保持生成视频的时序一致性。为解耦身份与姿态特征,构建了包含三张图像的属性-身份解耦三元组(AIDT)数据集,其中两图同姿态、两图同身份。结合全面的遮挡增强,提升了对遮挡的鲁棒性。此外,通过3D重建技术作为输入条件,增强对大姿态变化的处理能力。大量实验表明,本框架在身份保留、时序一致性和视觉质量方面均优于现有方法,且所需推理步骤更少。有效缓解了视频人脸替换中的时序闪烁、身份丢失、遮挡与姿态变化等关键挑战。
原文摘要 · Abstract (English)
Video face swapping is becoming increasingly popular across various applications, yet existing methods primarily focus on static images and struggle with video face swapping because of temporal consistency and complex scenarios. In this paper, we present the first diffusion-based framework specifically designed for video face swapping. Our approach introduces a novel image-video hybrid training framework that leverages both abundant static image data and temporal video sequences, addressing the inherent limitations of video-only training. The framework incorporates a specially designed diffusion model coupled with a VidFaceVAE that effectively processes both types of data to better maintain temporal coherence of the generated videos. To further disentangle identity and pose features, we construct the Attribute-Identity Disentanglement Triplet (AIDT) Dataset, where each triplet has three face images, with two images sharing the same pose and two sharing the same identity. Enhanced with a comprehensive occlusion augmentation, this dataset also improves robustness against occlusions. Additionally, we integrate 3D reconstruction techniques as input conditioning to our network for handling large pose variations. Extensive experiments demonstrate that our framework achieves superior performance in identity preservation, temporal consistency, and visual quality compared to existing methods, while requiring fewer inference steps. Our approach effectively mitigates key challenges in video face swapping, including temporal flickering, identity preservation, and robustness to occlusions and pose variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。