无需掩码的视频换头技术,提升表情与动作一致性。
DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping
- 用生成模型构建配对数据,实现无掩码换头训练
- 在真实视频上达到顶尖的图像质量与表情连贯性
- 适合视频编辑、数字人创作等场景研究者
视频换头旨在将目标人物的整个头部(包括面部身份、头型和发型)替换为参考图像内容,同时保留身体、背景和运动动态。由于缺乏真实配对的换头数据,现有方法通常基于同一视频内跨帧配对进行训练,并依赖掩码修复以减少身份泄露。但该范式易产生边界伪影,且无法恢复被掩码遮挡的关键线索,如面部姿态、表情和运动动态。为此,我们设计一个视频编辑模型,生成新头部作为假换头输入,同时保持帧间同步的面部姿态与表情,从而构建首个跨身份配对数据集HeadSwapBench,支持训练( TrainNum{}视频)与评测( TestNum{}视频)。在此基础上,提出DirectSwap:一种无掩码的直接视频换头框架,将图像U-Net扩展为带运动模块与条件输入的视频扩散模型。进一步引入运动与表情感知重建(MEAR)损失,通过帧差幅度与人脸关键点邻近度重加权扩散损失,提升跨帧运动与表情一致性。大量实验表明,DirectSwap在多样化野外视频场景中均实现顶尖的视觉质量、身份保真度及运动与表情一致性。代码与数据集将公开发布。
原文摘要 · Abstract (English)
Video head swapping aims to replace the entire head of a video subject, including facial identity, head shape, and hairstyle, with that of a reference image, while preserving the target body, background, and motion dynamics. Due to the lack of ground-truth paired swapping data, prior methods typically train on cross-frame pairs of the same person within a video and rely on mask-based inpainting to mitigate identity leakage. Beyond potential boundary artifacts, this paradigm struggles to recover essential cues occluded by the mask, such as facial pose, expressions, and motion dynamics. To address these issues, we prompt a video editing model to synthesize new heads for existing videos as fake swapping inputs, while maintaining frame-synchronized facial poses and expressions. This yields HeadSwapBench, the first cross-identity paired dataset for video head swapping, which supports both training (\TrainNum{} videos) and benchmarking (\TestNum{} videos) with genuine outputs. Leveraging this paired supervision, we propose DirectSwap, a mask-free, direct video head-swapping framework that extends an image U-Net into a video diffusion model with a motion module and conditioning inputs. Furthermore, we introduce the Motion- and Expression-Aware Reconstruction (MEAR) loss, which reweights the diffusion loss per pixel using frame-difference magnitudes and facial-landmark proximity, thereby enhancing cross-frame coherence in motion and expressions. Extensive experiments demonstrate that DirectSwap achieves state-of-the-art visual quality, identity fidelity, and motion and expression consistency across diverse in-the-wild video scenes. We will release the source code and the HeadSwapBench dataset to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。