arXiv:2412.21042cs.CVcs.MM2024-12被引 58

用扩散模型生成视觉风格提示,提升盲源人脸修复质量

Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration

  • 在生成模型隐空间中用扩散模型生成视觉提示
  • 新设计的风格调制聚合层增强特征提取能力
  • 适合需要高质量人脸修复的研究与应用

盲源人脸修复旨在从多种未知退化源中恢复高质量人脸图像,由于退化图像信息极少而极具挑战。基于先验知识的方法虽利用几何先验和面部特征取得进展,但难以捕捉精细细节。为此,我们提出一种视觉风格提示学习框架,利用扩散概率模型在预训练生成模型的隐空间中显式生成视觉提示,以指导修复过程。为充分挖掘视觉提示并增强信息丰富模式的提取,引入风格调制聚合变换层。大量实验与应用表明,该方法在实现高质量盲源人脸修复方面具有显著优势。代码已开源:https://github.com/LonglongaaaGo/VSPBFR。

原文摘要 · Abstract (English)

Blind face restoration aims to recover high-quality facial images from various unidentified sources of degradation, posing significant challenges due to the minimal information retrievable from the degraded images. Prior knowledge-based methods, leveraging geometric priors and facial features, have led to advancements in face restoration but often fall short of capturing fine details. To address this, we introduce a visual style prompt learning framework that utilizes diffusion probabilistic models to explicitly generate visual prompts within the latent space of pre-trained generative models. These prompts are designed to guide the restoration process. To fully utilize the visual prompts and enhance the extraction of informative and rich patterns, we introduce a style-modulated aggregation transformation layer. Extensive experiments and applications demonstrate the superiority of our method in achieving high-quality blind face restoration. The source code is available at \href{https://github.com/LonglongaaaGo/VSPBFR}{https://github.com/LonglongaaaGo/VSPBFR}.

人脸修复扩散模型风格提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。