用扩散模型生成人脸时保持身份一致,且逼真度更高。
Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models

- 构建21万张带身份标签的合成图像,提升训练时的身份感知。
- 在未见过的人脸上,身份相似度不输同类方法,但真实感更强。
- 提出新评估指标FIQ,同时衡量身份一致性和图像质量。
生成式扩散模型已革新人脸图像合成,但高分辨率输出中的身份一致性仍具挑战,尤其在安全系统与生物认证等隐私敏感场景中至关重要。本文提出Diff-ID,一种基于扩散模型的框架,在保证照片级真实感的同时强化身份一致性。核心是利用CelebA-HQ、FFHQ和LAION-Face合成并标注的21万张图像数据集,并通过微调的BLIP模型生成描述文本以增强身份意识。在微调后的Stable Diffusion UNet中,引入融合ArcFace与CLIP嵌入的双交叉注意力适配器;为进一步保障身份保真度,设计基于ArcFace余弦相似度的伪判别器损失,并采用指数时间步加权策略。在保留与未见人脸上的实验表明,Diff-ID在原始ArcFace人脸相似度上不逊于InstantID,但显著降低FID分数,并在身份-真实感权衡上表现最优。此外,提出统一的DDIM驱动变形管道,实现无需每身份微调的高质量面部插值。强调身份一致性与真实感应联合评估,提出互补性指标面像质量(FIQ),结合身份相似性与感知真实感,同时保留FS与FID作为主指标。
原文摘要 · Abstract (English)
Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is especially vital for security systems, biometric authentication, and privacy sensitive applications, where any drift in identity integrity can undermine trust and functionality. We introduce Diff-ID, a diffusion based framework that enforces identity consistency while delivering photorealistic quality. Central to our approach is a custom 210K image dataset synthesized from CelebA-HQ, FFHQ, and LAION-Face and captioned via a fine tuned BLIP model to bolster identity awareness during training. Diff-ID integrates ArcFace and CLIP embeddings through a dual cross attention adapter within a fine tuned Stable Diffusion UNet. To further reinforce identity fidelity, we propose a pseudo discriminator loss based on ArcFace cosine similarity with exponential timestep weighting. Experiments on held out and unseen faces show that Diff-ID does not exceed InstantID in raw ArcFace Face Similarity, but achieves substantially lower FID and the strongest FIQ based identity--realism trade off among the evaluated methods. We also present a unified DDIM based morphing pipeline that enables qualitative facial interpolation without per identity fine tuning. We further argue that identity preservation and photorealism should be evaluated jointly rather than in isolation, as high identity similarity alone does not guarantee realistic outputs. To make this trade off explicit, we report Face Image Quality (FIQ) as a complementary ratio based score that combines identity similarity and perceptual realism while keeping FS and FID as the primary metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。