无需真实标签,用扩散模型自动生成修复目标,实现无监督盲脸修复。
Towards Unsupervised Blind Face Restoration using Diffusion Prior
- 利用预训练扩散模型生成符合自然图像分布的伪目标图。
- 在无真实标签情况下,显著提升模型对未知退化的修复效果。
- 适合缺乏标注数据但需高保真修复的场景,如真实世界人脸恢复。
盲脸修复方法在大规模合成数据集上表现优异,这些数据通常通过手工设计的退化流程生成。然而,基于合成退化的模型难以处理未见过的退化类型。本文提出仅使用一组未知退化且无真实标签的输入图像,微调一个可将输入映射为清晰、上下文一致输出的修复模型。我们利用预训练扩散模型作为生成先验,在保持输入内容一致性的约束下,从自然图像分布中生成高质量图像作为伪目标,用于微调预训练修复模型。与多数在测试时使用扩散模型的方法不同,我们的方法仅在训练阶段使用扩散模型,从而保证推理高效。大量实验表明,该方法能持续提升预训练盲脸修复模型的感知质量,并保持与输入内容的高度一致性。最佳模型在合成与真实世界数据集上均达到当前最优性能。
原文摘要 · Abstract (English)
Blind face restoration methods have shown remarkable performance, particularly when trained on large-scale synthetic datasets with supervised learning. These datasets are often generated by simulating low-quality face images with a handcrafted image degradation pipeline. The models trained on such synthetic degradations, however, cannot deal with inputs of unseen degradations. In this paper, we address this issue by using only a set of input images, with unknown degradations and without ground truth targets, to fine-tune a restoration model that learns to map them to clean and contextually consistent outputs. We utilize a pre-trained diffusion model as a generative prior through which we generate high quality images from the natural image distribution while maintaining the input image content through consistency constraints. These generated images are then used as pseudo targets to fine-tune a pre-trained restoration model. Unlike many recent approaches that employ diffusion models at test time, we only do so during training and thus maintain an efficient inference-time performance. Extensive experiments show that the proposed approach can consistently improve the perceptual quality of pre-trained blind face restoration models while maintaining great consistency with the input contents. Our best model also achieves the state-of-the-art results on both synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。