arXiv:2604.24493cs.CV2026-04

用扩散模型实现高保真人脸换脸,身份一致性更强。

CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping

论文配图:CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping
图 1 · 摘自论文原文
  • 通过多尺度交叉注意力融合姿态、表情与身份信息
  • FID达11.73,显著优于FaceShifter等基线方法
  • 支持跨姿态表情的精细区域控制,适合高精度人脸编辑

人脸换脸旨在将源人脸的身份特征迁移到目标人脸,同时保持姿态、表情和上下文一致性。现有基于GAN的方法常因可控性差和模式崩溃,难以平衡身份保留与视觉真实感。本文提出首个基于扩散模型的人脸换脸方法CA-IDD,通过多尺度交叉注意力整合注视方向、身份嵌入和面部解析的多模态引导。预计算的身份嵌入通过分层注意力层融入去噪过程,实现精准一致的身份迁移。为提升语义连贯性与视觉质量,引入专家指导监督,包括面部解析与注视一致性模块。相比GAN或隐式融合方法,该扩散框架具备稳定训练、强泛化能力及空间自适应身份对齐,支持跨姿态和表情变化的细粒度区域控制。实验显示,CA-IDD达到FID 11.73,优于FaceShifter和MegaFS等基线。定性结果也表明其在多种姿态下身份保留更优,为未来扩散模型人脸编辑奠定坚实基础。

原文摘要 · Abstract (English)

Face swapping aims to optimize realistic facial image generation by leveraging the identity of a source face onto a target face while preserving pose, expression, and context. However, existing methods, especially GAN-based methods, often struggle to balance identity preservation and visual realism due to limited controllability and mode collapse. In this paper, we introduce CA-IDD (Cross-Attention Guided Identity-Conditional Diffusion), the first diffusion-based face swapping approach that integrates multi-modal guidance comprising gaze, identity, and facial parsing through multi-scale cross-attention. Precomputed identity embeddings are incorporated into the denoising process via hierarchical attention layers, resulting in accurate and consistent identity transfer. To improve semantic coherence and visual quality, we use expert-guided supervision, with facial parsing and gaze-consistency modules. Unlike GAN-based or implicit-fusion methods, our diffusion framework provides stable training, robust generalization, and spatially adaptive identity alignment, allowing for fine-grained regional control across pose and expression variations. CA-IDD achieves an FID of 11.73, exceeding established baselines such as FaceShifter and MegaFS. Qualitative results also reveal improved identity retention across diverse poses, establishing CA-IDD as a strong foundation for future diffusion-based face editing.

人脸换脸扩散模型身份一致图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。