arXiv:2501.06546cs.CVcs.AI2025-01被引 3

用文字描述指导暗光图像增强,提升视觉效果与指标一致性。

Natural Language Supervision for Low-light Image Enhancement

  • 引入文本监督机制,通过文字描述对齐图像与语义特征。
  • 设计文本引导条件机制,捕捉跨模态细粒度关联,提升生成质量。
  • 适合关注多模态学习、图像增强与语义对齐的研究者。

随着深度学习发展,低光照图像增强(LLIE)方法取得了显著进展。主流方法依赖成对的暗光与正常光照图像进行端到端映射学习,但不同光照下的正常光照图像难以定义‘完美’参考,导致指标优化与视觉感知难以兼顾。受跨模态研究启发,本文提出自然语言监督(NLS)策略,利用图像对应的文本特征图作为指导,提供通用且灵活的图像描述接口。由于文本描述条件下的图像分布高度多模态,训练困难,为此设计了文本引导条件机制(TCM),建模图像区域与句子词汇间的关联,增强细粒度跨模态对齐能力。进一步设计信息融合注意力(IFA)模块,有效整合多层级图像与文本特征。将TCM与IFA集成至名为NaLSuper的网络中,实验表明该方法在多个数据集上具有更强鲁棒性与优越性能。

原文摘要 · Abstract (English)

With the development of deep learning, numerous methods for low-light image enhancement (LLIE) have demonstrated remarkable performance. Mainstream LLIE methods typically learn an end-to-end mapping based on pairs of low-light and normal-light images. However, normal-light images under varying illumination conditions serve as reference images, making it difficult to define a ``perfect'' reference image This leads to the challenge of reconciling metric-oriented and visual-friendly results. Recently, many cross-modal studies have found that side information from other related modalities can guide visual representation learning. Based on this, we introduce a Natural Language Supervision (NLS) strategy, which learns feature maps from text corresponding to images, offering a general and flexible interface for describing an image under different illumination. However, image distributions conditioned on textual descriptions are highly multimodal, which makes training difficult. To address this issue, we design a Textual Guidance Conditioning Mechanism (TCM) that incorporates the connections between image regions and sentence words, enhancing the ability to capture fine-grained cross-modal cues for images and text. This strategy not only utilizes a wider range of supervised sources, but also provides a new paradigm for LLIE based on visual and textual feature alignment. In order to effectively identify and merge features from various levels of image and textual information, we design an Information Fusion Attention (IFA) module to enhance different regions at different levels. We integrate the proposed TCM and IFA into a Natural Language Supervision network for LLIE, named NaLSuper. Finally, extensive experiments demonstrate the robustness and superior effectiveness of our proposed NaLSuper.

图像增强多模态文本引导低光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。