arXiv:2602.07835cs.CV2026-02中稿 · WACV 2026

无需训练即可实现高质量视频换脸,保持人脸一致性与时间连贯性。

VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping

  • 通过频谱注意力插值保留关键身份特征。
  • 采用目标结构引导增强帧间结构对齐效果。
  • 流引导注意力时序平滑,提升生成连贯性。

我们提出一种无需训练、即插即用的视频人脸换脸方法VFace,可无缝集成于基于扩散模型的图像级换脸框架。首先,引入频谱注意力插值技术,以促进生成并保持关键身份特征的完整性;其次,通过即插即用的注意力注入实现目标结构引导,更好对齐目标帧的结构特征;第三,提出流引导注意力时序平滑机制,在不修改底层扩散模型的前提下,增强时空一致性,缓解逐帧生成带来的时序不一致问题。该方法无需额外训练或视频特定微调。大量实验表明,本方法显著提升了时间一致性与视觉保真度,为视频人脸换脸提供了一种实用且模块化的新方案。代码已开源:https://github.com/Sanoojan/VFace。

原文摘要 · Abstract (English)

We present a training-free, plug-and-play method, namely VFace, for high-quality face swapping in videos. It can be seamlessly integrated with image-based face swapping approaches built on diffusion models. First, we introduce a Frequency Spectrum Attention Interpolation technique to facilitate generation and intact key identity characteristics. Second, we achieve Target Structure Guidance via plug-and-play attention injection to better align the structural features from the target frame to the generation. Third, we present a Flow-Guided Attention Temporal Smoothening mechanism that enforces spatiotemporal coherence without modifying the underlying diffusion model to reduce temporal inconsistencies typically encountered in frame-wise generation. Our method requires no additional training or video-specific fine-tuning. Extensive experiments show that our method significantly enhances temporal consistency and visual fidelity, offering a practical and modular solution for video-based face swapping. Our code is available at https://github.com/Sanoojan/VFace.

视频换脸扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。