针对超大倍数模糊人脸,用身份解耦与适配提升还原真实度。
Personalized Face Super-Resolution with Identity Decoupling and Fitting
- 掩码低分辨率人脸区域,去除不可靠身份线索。
- 用参考图对齐提供风格引导,结合真值提取的身份嵌入精细建模。
- 轻量微调身份嵌入,显著提升大倍率下的身份一致性和视觉质量。
近年来,人脸超分辨率(FSR)方法在标准条件下已取得显著进展,能保持高图像保真度和身份一致性。然而,在极端退化场景(如缩放倍数>8×)下,输入图像中的关键属性和身份信息常严重丢失,传统模型难以重建出真实且身份一致的人脸,易产生幻觉结果。为此,我们提出身份解耦与适配的新型FSR方法(IDFSR),旨在提升大缩放倍数下的身份恢复能力并减少幻觉。方法包含三项设计:1)掩码低分辨率(LR)图像中的人脸区域,消除不可靠身份线索;2)将参考图像进行形变对齐以匹配LR输入,提供风格指导;3)利用真值图像提取的身份嵌入进行细粒度身份建模与个性化适配。首先预训练一个基于扩散模型的方法,通过强制其使用风格和身份嵌入重建掩码后的LR人脸区域,显式解耦风格与身份。随后冻结大部分网络参数,仅对少量目标身份图像进行轻量级身份嵌入微调。该嵌入编码了细粒度面部特征与精确身份信息,显著提升了身份一致性和感知质量。大量定量评估与视觉对比表明,所提方法在极端退化条件下明显优于现有方法,尤其在身份一致性上表现优异。
原文摘要 · Abstract (English)
In recent years, face super-resolution (FSR) methods have achieved remarkable progress, generally maintaining high image fidelity and identity (ID) consistency under standard settings. However, in extreme degradation scenarios (e.g., scale $> 8\times$), critical attributes and ID information are often severely lost in the input image, making it difficult for conventional models to reconstruct realistic and ID-consistent faces. Existing methods tend to generate hallucinated faces under such conditions, producing restored images lacking authentic ID constraints. To address this challenge, we propose a novel FSR method with Identity Decoupling and Fitting (IDFSR), designed to enhance ID restoration under large scaling factors while mitigating hallucination effects. Our approach involves three key designs: 1) \textbf{Masking} the facial region in the low-resolution (LR) image to eliminate unreliable ID cues; 2) \textbf{Warping} a reference image to align with the LR input, providing style guidance; 3) Leveraging \textbf{ID embeddings} extracted from ground truth (GT) images for fine-grained ID modeling and personalized adaptation. We first pretrain a diffusion-based model to explicitly decouple style and ID by forcing it to reconstruct masked LR face regions using both style and identity embeddings. Subsequently, we freeze most network parameters and perform lightweight fine-tuning of the ID embedding using a small set of target ID images. This embedding encodes fine-grained facial attributes and precise ID information, significantly improving both ID consistency and perceptual quality. Extensive quantitative evaluations and visual comparisons demonstrate that the proposed IDFSR substantially outperforms existing approaches under extreme degradation, particularly achieving superior performance on ID consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。