arXiv:2410.04161cs.CV2024-10ICLR被引 15

用多模态信息修复低质人脸,避免生成错误特征和身份

Overcoming False Illusions in Real-World Face Restoration with Multi-Modal Guided Diffusion Model

  • 结合文本提示、高清参考图和身份信息进行联合引导
  • 在严重退化图像上恢复出细节更真实的人脸,身份保持准确
  • 适合需要高保真人脸修复与身份一致性的实际应用

我们提出一种新型多模态引导的真实世界人脸修复方法(MGFR),旨在提升从低质量输入中恢复人脸的品质。该方法融合属性文本提示、高质量参考图像和身份信息,有效缓解生成错误面部特征和身份的问题。通过引入双控制适配器与两阶段训练策略,充分利用多模态先验信息完成定向修复任务。同时,我们构建了Reface-HQ数据集,包含4800个身份的超过21,000张高分辨率人脸图像,以支持参考图像训练需求。实验表明,该方法在严重退化条件下仍能实现优异的视觉质量,显著提升身份保真度与属性修正能力。训练中加入负样本及属性提示进一步增强了模型生成细节丰富且感知真实的图像能力。

原文摘要 · Abstract (English)

We introduce a novel Multi-modal Guided Real-World Face Restoration (MGFR) technique designed to improve the quality of facial image restoration from low-quality inputs. Leveraging a blend of attribute text prompts, high-quality reference images, and identity information, MGFR can mitigate the generation of false facial attributes and identities often associated with generative face restoration methods. By incorporating a dual-control adapter and a two-stage training strategy, our method effectively utilizes multi-modal prior information for targeted restoration tasks. We also present the Reface-HQ dataset, comprising over 21,000 high-resolution facial images across 4800 identities, to address the need for reference face training images. Our approach achieves superior visual quality in restoring facial details under severe degradation and allows for controlled restoration processes, enhancing the accuracy of identity preservation and attribute correction. Including negative quality samples and attribute prompts in the training further refines the model's ability to generate detailed and perceptually accurate images.

人脸修复扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。