arXiv:2504.19127cs.CVcs.MM2025-04中稿 · ICMR 2025 Main tra…被引 19

用多模态语义先验提升暗光图像增强效果

DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning

  • 结合图像与文本语义先验,指导暗光图像增强
  • 在五个数据集上优于当前最优方法,显著改善暗区细节
  • 适合需要高质量图像增强的应用场景

长期以来,高层语义学习被认为能促进多种下游计算机视觉任务。然而,在暗光图像增强(LLIE)领域,现有方法仅建立暗光与正常光照域之间的粗略映射,未考虑不同区域的语义信息,尤其在严重信息丢失的极暗区域表现不佳。为此,本文提出基于Retinex分解的深度语义先验引导框架(DeepSPG),利用预训练语义分割模型和多模态学习探索有信息量的语义知识。关键创新在于融合图像级与文本级语义先验,构建组合式深度语义先验引导的多模态学习框架:通过层次化语义特征提供图像级先验;借助预训练视觉-语言模型引入自然语言语义约束作为文本级先验;设计多尺度语义感知结构以高效融合语义特征。最终,DeepSPG在五个基准数据集上均超越现有最优方法,性能显著提升。代码与实现细节已公开于https://github.com/Wenyuzhy/DeepSPG。

原文摘要 · Abstract (English)

There has long been a belief that high-level semantics learning can benefit various downstream computer vision tasks. However, in the low-light image enhancement (LLIE) community, existing methods learn a brutal mapping between low-light and normal-light domains without considering the semantic information of different regions, especially in those extremely dark regions that suffer from severe information loss. To address this issue, we propose a new deep semantic prior-guided framework (DeepSPG) based on Retinex image decomposition for LLIE to explore informative semantic knowledge via a pre-trained semantic segmentation model and multimodal learning. Notably, we incorporate both image-level semantic prior and text-level semantic prior and thus formulate a multimodal learning framework with combinatorial deep semantic prior guidance for LLIE. Specifically, we incorporate semantic knowledge to guide the enhancement process via three designs: an image-level semantic prior guidance by leveraging hierarchical semantic features from a pre-trained semantic segmentation model; a text-level semantic prior guidance by integrating natural language semantic constraints via a pre-trained vision-language model; a multi-scale semantic-aware structure that facilitates effective semantic feature incorporation. Eventually, our proposed DeepSPG demonstrates superior performance compared to state-of-the-art methods across five benchmark datasets. The implementation details and code are publicly available at https://github.com/Wenyuzhy/DeepSPG.

图像增强语义先验多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。