arXiv:2412.20916cs.CV2024-12中稿 · AAAI被引 43

用视觉语言模型生成感知先验,提升暗光图像真实感增强效果

Low-Light Image Enhancement via Generative Perceptual Priors

  • 基于视觉语言模型提取全局与局部感知先验指导增强
  • 在真实场景数据上优于现有最先进方法,泛化能力更强
  • 适合需要高真实感暗光图像增强的科研与工业应用

尽管暗光图像增强在提升可视性、恢复纹理细节和抑制噪声方面取得显著进展,但现有方法在真实场景中的应用仍受限于光照条件多样性。同时,生成视觉上真实且吸引人的结果仍是未充分探索的方向。为此,我们提出一种新型低光照图像增强框架(GPP-LLIE),利用视觉语言模型生成生成式感知先验。首先设计一个流程,引导视觉语言模型评估暗光图像的多个视觉属性,并量化评估结果以输出全局与局部感知先验。随后,为将这些先验融入增强过程,我们在扩散模型中引入基于Transformer的主干网络,设计了受全局与局部感知先验引导的新层归一化(GPP-LN)和注意力机制(LPP-Attn)。大量实验表明,该模型在配对暗光数据集上超越当前最先进方法,并在真实世界数据上展现出更优泛化性能。代码已公开于 https://github.com/LowLevelAI/GPP-LLIE。

原文摘要 · Abstract (English)

Although significant progress has been made in enhancing visibility, retrieving texture details, and mitigating noise in Low-Light (LL) images, the challenge persists in applying current Low-Light Image Enhancement (LLIE) methods to real-world scenarios, primarily due to the diverse illumination conditions encountered. Furthermore, the quest for generating enhancements that are visually realistic and attractive remains an underexplored realm. In response to these challenges, we introduce a novel \textbf{LLIE} framework with the guidance of \textbf{G}enerative \textbf{P}erceptual \textbf{P}riors (\textbf{GPP-LLIE}) derived from vision-language models (VLMs). Specifically, we first propose a pipeline that guides VLMs to assess multiple visual attributes of the LL image and quantify the assessment to output the global and local perceptual priors. Subsequently, to incorporate these generative perceptual priors to benefit LLIE, we introduce a transformer-based backbone in the diffusion process, and develop a new layer normalization (\textit{\textbf{GPP-LN}}) and an attention mechanism (\textit{\textbf{LPP-Attn}}) guided by global and local perceptual priors. Extensive experiments demonstrate that our model outperforms current SOTA methods on paired LL datasets and exhibits superior generalization on real-world data. The code is released at \url{https://github.com/LowLevelAI/GPP-LLIE}.

暗光增强视觉语言模型扩散模型感知先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。