解决人脸修复中大遮挡下的身份保持难题
When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions

- 多参考图像提取身份特征,构建语义先验指导修复
- 在严重遮挡下身份保留率提升,文本控制修复更精准
- 适合需要高保真人脸重建的图像编辑场景
基于扩散模型的人脸修复虽已取得卓越视觉效果,但在严重遮挡与冲突文本引导下仍难以保持身份一致性。为此,我们提出参考语义人脸修复(ReSem-Face),一种级联扩散框架,引入显式身份条件语义先验,实现多参考图像下的人脸修复。该方法从多个参考图像中提炼代表性身份特征,用于重建缺失语义区域,并通过多流条件架构引导扩散过程。此设计在像素缺失时提供强语义约束,稳定身份重建,同时兼容提示词驱动的编辑。在CelebAHQ-IDI-5和VGGFace2数据集上的实验表明,相较于主流基线,ReSem-Face在严重语义掩码下能更可靠地保持身份一致性,并提升文本控制修复质量。
原文摘要 · Abstract (English)
Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascaded diffusion framework that introduces an explicit identity-conditioned semantic prior for multi-reference face inpainting. Our approach distills representative identity features from multiple references to reconstruct missing semantic regions, which then guide the diffusion process through a multi-stream conditioning architecture. This design provides strong semantic constraints when pixels are absent and stabilizes identity reconstruction while remaining compatible with prompt-driven edits. Experiments on CelebAHQ-IDI-5 and VGGFace2 demonstrate that ReSem-Face yields more reliable identity-preserving completion under severe semantic masks and improves text-controlled editing quality compared with representative baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。