DreamID通过三元组身份分组实现高保真人脸替换,0.6秒完成512分辨率生成。
DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning
- 构建三元组身份分组数据,显式监督提升身份相似度与属性保持
- 结合SD Turbo加速扩散模型,单步推理实现端到端像素级训练
- 支持眼镜、脸型等特定属性精细调整,适用于复杂光照与遮挡场景
本文提出DreamID,一种基于扩散模型的人脸替换方法,在身份相似性、属性保持、图像保真度和推理速度方面均达到高水平。不同于传统依赖隐式监督的训练方式,DreamID通过构建三元组身份分组数据,实现显式监督,显著提升身份一致性和属性保留能力。扩散模型的迭代特性使使用高效图像空间损失函数困难,因训练中需耗时多步采样生成图像。为此,我们采用加速扩散模型SD Turbo,将推理步骤缩减至单次迭代,实现高效的像素级端到端训练。此外,提出包含SwapNet、FaceNet和ID Adapter的改进架构,充分释放三元组身份分组显式监督的优势。进一步地,训练中显式修改三元组身份分组数据,可对眼镜、脸型等特定属性进行微调与保留。大量实验表明,DreamID在身份相似性、姿态与表情保持及图像保真度上优于现有方法。整体上,该方法在512×512分辨率下仅需0.6秒即可生成高质量人脸替换结果,并在复杂光照、大角度与遮挡等挑战性场景中表现优异。
原文摘要 · Abstract (English)
In this paper, we introduce DreamID, a diffusion-based face swapping model that achieves high levels of ID similarity, attribute preservation, image fidelity, and fast inference speed. Unlike the typical face swapping training process, which often relies on implicit supervision and struggles to achieve satisfactory results. DreamID establishes explicit supervision for face swapping by constructing Triplet ID Group data, significantly enhancing identity similarity and attribute preservation. The iterative nature of diffusion models poses challenges for utilizing efficient image-space loss functions, as performing time-consuming multi-step sampling to obtain the generated image during training is impractical. To address this issue, we leverage the accelerated diffusion model SD Turbo, reducing the inference steps to a single iteration, enabling efficient pixel-level end-to-end training with explicit Triplet ID Group supervision. Additionally, we propose an improved diffusion-based model architecture comprising SwapNet, FaceNet, and ID Adapter. This robust architecture fully unlocks the power of the Triplet ID Group explicit supervision. Finally, to further extend our method, we explicitly modify the Triplet ID Group data during training to fine-tune and preserve specific attributes, such as glasses and face shape. Extensive experiments demonstrate that DreamID outperforms state-of-the-art methods in terms of identity similarity, pose and expression preservation, and image fidelity. Overall, DreamID achieves high-quality face swapping results at 512*512 resolution in just 0.6 seconds and performs exceptionally well in challenging scenarios such as complex lighting, large angles, and occlusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。