arXiv:2410.09864cs.CV2024-10被引 6

用专业摄影图训练扩散模型,让修复人脸更真实

AuthFace: Towards Authentic Blind Face Restoration with Face-oriented Generative Diffusion Prior

  • 用1.5千张8K以上高清人像微调扩散模型,聚焦面部特征
  • 在真实数据集上显著提升眼、嘴等关键部位修复质量
  • 适合需要高保真人脸修复的图像增强与数字内容创作

盲人脸修复(BFR)是计算机视觉中的基础且挑战性问题。现有方法多依赖预训练文本到图像(T2I)扩散模型的人脸先验,但常生成非人脸特征且面部细节不足,难以实用。本文提出AuthFace框架,通过面向人脸的生成扩散先验实现高度真实的修复。首先收集1.5K张由专业摄影师拍摄的超8K分辨率高质量图像;基于此,设计新的修复微调流程,结合质量优先与摄影引导标注,由摄影师参与润色与审核,充分挖掘高质量照片潜力;由此可微妙利用预训练模型的自然图像先验,显著增强其面部细节恢复能力。此外,为减少眼部、口部等关键区域伪影,提出时间感知潜在面部特征损失,学习更真实的修复过程。在合成与真实世界BFR数据集上的大量实验表明本方法优越。

原文摘要 · Abstract (English)

Blind face restoration (BFR) is a fundamental and challenging problem in computer vision. To faithfully restore high-quality (HQ) photos from poor-quality ones, recent research endeavors predominantly rely on facial image priors from the powerful pretrained text-to-image (T2I) diffusion models. However, such priors often lead to the incorrect generation of non-facial features and insufficient facial details, thus rendering them less practical for real-world applications. In this paper, we propose a novel framework, namely AuthFace that achieves highly authentic face restoration results by exploring a face-oriented generative diffusion prior. To learn such a prior, we first collect a dataset of 1.5K high-quality images, with resolutions exceeding 8K, captured by professional photographers. Based on the dataset, we then introduce a novel face-oriented restoration-tuning pipeline that fine-tunes a pretrained T2I model. Identifying key criteria of quality-first and photography-guided annotation, we involve the retouching and reviewing process under the guidance of photographers for high-quality images that show rich facial features. The photography-guided annotation system fully explores the potential of these high-quality photographic images. In this way, the potent natural image priors from pretrained T2I diffusion models can be subtly harnessed, specifically enhancing their capability in facial detail restoration. Moreover, to minimize artifacts in critical facial areas, such as eyes and mouth, we propose a time-aware latent facial feature loss to learn the authentic face restoration process. Extensive experiments on the synthetic and real-world BFR datasets demonstrate the superiority of our approach.

人脸修复扩散模型图像增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。