arXiv:2502.02215cs.CV2025-02ICLR被引 13

用低质图像作为潜在一致性模型中间态,实现高效高保真盲人脸修复。

InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration

  • 将低质图视为潜在一致性模型的中间状态,提升修复语义一致性。
  • 在合成与真实数据集上均超越现有方法,且推理速度更快。
  • 融合感知损失与空间特征提取,适合真实场景下的人脸修复任务。

扩散先验被用于盲人脸修复(BFR),通过在修复数据集上微调扩散模型(DMs)来恢复低质量图像。然而,直接应用DMs存在若干关键局限:(i) 扩散先验语义一致性差(如身份、结构和色彩),增加模型优化难度;(ii) 依赖数百次去噪迭代,难以有效结合感知损失,而后者对忠实还原至关重要。观察到潜在一致性模型(LCM)在常微分方程轨迹上学习噪声到数据的一致性映射,因而表现出更强的主体身份、结构信息和色彩保留能力,我们提出InterLCM,利用LCM的优异语义一致性和效率来克服上述问题。将低质量图像视为LCM的中间状态,InterLCM通过从较早的LCM步骤开始,在保真度与质量间取得平衡。LCM还支持训练中引入感知损失,显著提升真实场景下的修复质量。为缓解结构与语义不确定性,InterLCM引入视觉模块提取视觉特征,以及空间编码器捕捉空间细节,增强重建图像保真度。大量实验表明,InterLCM在合成与真实数据集上均优于现有方法,并实现更快的推理速度。

原文摘要 · Abstract (English)

Diffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations. (i) The diffusion prior has inferior semantic consistency (e.g., ID, structure and color.), increasing the difficulty of optimizing the BFR model; (ii) reliance on hundreds of denoising iterations, preventing the effective cooperation with perceptual losses, which is crucial for faithful restoration. Observing that the latent consistency model (LCM) learns consistency noise-to-data mappings on the ODE-trajectory and therefore shows more semantic consistency in the subject identity, structural information and color preservation, we propose InterLCM to leverage the LCM for its superior semantic consistency and efficiency to counter the above issues. Treating low-quality images as the intermediate state of LCM, InterLCM achieves a balance between fidelity and quality by starting from earlier LCM steps. LCM also allows the integration of perceptual loss during training, leading to improved restoration quality, particularly in real-world scenarios. To mitigate structural and semantic uncertainties, InterLCM incorporates a Visual Module to extract visual features and a Spatial Encoder to capture spatial details, enhancing the fidelity of restored images. Extensive experiments demonstrate that InterLCM outperforms existing approaches in both synthetic and real-world datasets while also achieving faster inference speed.

人脸修复扩散模型低质图像LCM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。