arXiv:2512.07951cs.CV2025-12中稿 · CVPR

用关键帧引导实现电影级高清人脸替换,保持表情光影一致

Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality

  • 以关键帧为条件,结合视频参考指导身份融合
  • 在长视频中实现高保真重建与稳定身份保留
  • 自建数据集支持可靠训练,降低制作人工成本

视频人脸替换在影视制作中至关重要,但在长而复杂的视频序列中实现高保真度与时间一致性仍是重大挑战。受参考引导图像编辑进展启发,本文探索是否可利用源视频中的丰富视觉属性来提升人脸替换的保真度与时序连贯性。基于此,提出首个视频参考引导的人脸替换模型LivingSwap。该方法以关键帧作为条件信号注入目标身份,实现灵活可控的编辑。通过结合关键帧条件与视频参考引导,模型完成时序拼接,确保长视频序列中身份稳定与高保真重建。为解决参考引导训练数据稀缺问题,构建了配对的人脸替换数据集Face2Face,并反向数据对以保证可靠的监督信号。大量实验表明,该方法能无缝融合目标身份与源视频的表情、光照与运动特征,显著减少制作流程中的人工干预。

原文摘要 · Abstract (English)

Video face swapping is crucial in film and entertainment production, where achieving high fidelity and temporal consistency over long and complex video sequences remains a significant challenge. Inspired by recent advances in reference-guided image editing, we explore whether rich visual attributes from source videos can be similarly leveraged to enhance both fidelity and temporal coherence in video face swapping. Building on this insight, this work presents LivingSwap, the first video reference guided face swapping model. Our approach employs keyframes as conditioning signals to inject the target identity, enabling flexible and controllable editing. By combining keyframe conditioning with video reference guidance, the model performs temporal stitching to ensure stable identity preservation and high-fidelity reconstruction across long video sequences. To address the scarcity of data for reference-guided training, we construct a paired face-swapping dataset, Face2Face, and further reverse the data pairs to ensure reliable ground-truth supervision. Extensive experiments demonstrate that our method achieves state-of-the-art results, seamlessly integrating the target identity with the source video's expressions, lighting, and motion, while significantly reducing manual effort in production workflows. Project webpage: https://aim-uofa.github.io/LivingSwap

人脸替换视频生成参考引导电影制作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。