用参考图提升人脸修复真实度,避免生成失真
ReF-LDM: A Latent Diffusion Model for Reference-based Face Image Restoration
- 引入多张高清参考图,结合低清输入生成更真实人脸
- 通过缓存机制和时序加权损失,精准保留人脸特征
- 构建2万+对参考数据集,适合需要高保真修复的研究者
尽管近期盲人脸修复方法已能从低质量(LQ)输入生成细节丰富的高质量(HQ)图像,但生成内容可能偏离真实人物外貌。为解决此问题,引入高质量个人图像作为额外参考输入是一种有前景的策略。受潜在扩散模型(LDM)成功的启发,我们提出ReF-LDM,一种基于LDM的改进模型,可同时利用一张LQ图像与多张HQ参考图像生成高质量人脸图像。该模型集成了一种高效机制CacheKV,用于在生成过程中有效利用参考图像信息;此外,设计了时序缩放的身份损失,使模型更专注于学习人脸的判别性特征。最后,我们构建了FFHQ-Ref数据集,包含20,405对高质量人脸图像及其对应参考图像,可作为参考式人脸修复模型的训练与评估基准。
原文摘要 · Abstract (English)
While recent works on blind face image restoration have successfully produced impressive high-quality (HQ) images with abundant details from low-quality (LQ) input images, the generated content may not accurately reflect the real appearance of a person. To address this problem, incorporating well-shot personal images as additional reference inputs could be a promising strategy. Inspired by the recent success of the Latent Diffusion Model (LDM), we propose ReF-LDM, an adaptation of LDM designed to generate HQ face images conditioned on one LQ image and multiple HQ reference images. Our model integrates an effective and efficient mechanism, CacheKV, to leverage the reference images during the generation process. Additionally, we design a timestep-scaled identity loss, enabling our LDM-based model to focus on learning the discriminating features of human faces. Lastly, we construct FFHQ-Ref, a dataset consisting of 20,405 high-quality (HQ) face images with corresponding reference images, which can serve as both training and evaluation data for reference-based face restoration models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。