将图像级人脸替换优势迁移至视频,实现高保真、连贯的人脸替换。
DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
- 构建双向身份四元组数据,用扩散变换器显式监督身份注入。
- 在复杂场景下保持身份一致性和视觉真实感,优于现有方法。
- 适合需要高精度人脸替换的应用,如影视制作与虚拟演播。
视频人脸替换(VFS)需在不破坏原视频姿态、表情、光照、背景及动态信息的前提下,无缝融合源身份。现有方法难以兼顾身份相似性与属性保持,且缺乏时间一致性。为此,本文提出完整框架,将图像级人脸替换(IFS)的优势迁移至视频领域。首先设计同步身份数据流水线 SyncID-Pipe,预训练身份锚定视频生成器,并与 IFS 模型结合构建双向身份四元组以实现显式监督。基于此数据,提出首个基于扩散变换器的 DreamID-V 框架,核心为模态感知条件模块,可区分性注入多模态条件。同时引入合成到真实课程机制与身份一致性强化学习策略,提升复杂场景下的视觉真实感与身份连贯性。为解决基准匮乏问题,构建涵盖多样场景的 IDBench-V 综合评测基准。大量实验表明,DreamID-V 显著超越现有最优方法,展现出卓越泛化能力,可无缝适配多种替换任务。
原文摘要 · Abstract (English)
Video Face Swapping (VFS) requires seamlessly injecting a source identity into a target video while meticulously preserving the original pose, expression, lighting, background, and dynamic information. Existing methods struggle to maintain identity similarity and attribute preservation while preserving temporal consistency. To address the challenge, we propose a comprehensive framework to seamlessly transfer the superiority of Image Face Swapping (IFS) to the video domain. We first introduce a novel data pipeline SyncID-Pipe that pre-trains an Identity-Anchored Video Synthesizer and combines it with IFS models to construct bidirectional ID quadruplets for explicit supervision. Building upon paired data, we propose the first Diffusion Transformer-based framework DreamID-V, employing a core Modality-Aware Conditioning module to discriminatively inject multi-model conditions. Meanwhile, we propose a Synthetic-to-Real Curriculum mechanism and an Identity-Coherence Reinforcement Learning strategy to enhance visual realism and identity consistency under challenging scenarios. To address the issue of limited benchmarks, we introduce IDBench-V, a comprehensive benchmark encompassing diverse scenes. Extensive experiments demonstrate DreamID-V outperforms state-of-the-art methods and further exhibits exceptional versatility, which can be seamlessly adapted to various swap-related tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。