arXiv:2412.07693cs.CVeess.IV2024-12中稿 · the IEEE Transacti…被引 6

用CLIP模型提升低光图像增强效果,让机器看得更准。

Leveraging Content and Context Cues for Low-Light Image Enhancement

  • 通过提示学习构建无标注数据的图像先验
  • 融合内容与上下文线索,减少过饱和和噪声放大
  • 在分类、检测任务中验证有效,适合下游视觉模型

低光条件严重影响机器认知,制约计算机视觉系统在真实场景中的表现。由于低光数据稀缺且难以标注,本文聚焦于图像处理而非模型微调,以提升下游任务性能。提出利用CLIP模型捕捉图像先验并提供语义引导:首先设计一种数据增强策略,基于图像采样进行提示学习,无需成对或非成对正常光照数据即可学习图像先验;其次提出语义引导策略,充分利用现有低光标注数据,引入图像训练块的内容与上下文线索。定性实验表明,该方法显著提升整体对比度与色相,改善前景背景区分度,减少过饱和和噪声过度放大,这是零参考方法的常见问题。为验证对机器认知的有效性,不依赖人眼感知与任务性能的相关性假设,本文在多个低光数据集上进行消融研究,并对比相关零参考方法在图像分类、目标检测与人脸检测任务中的表现,证明了所提方法的优越性。

原文摘要 · Abstract (English)

Low-light conditions have an adverse impact on machine cognition, limiting the performance of computer vision systems in real life. Since low-light data is limited and difficult to annotate, we focus on image processing to enhance low-light images and improve the performance of any downstream task model, instead of fine-tuning each of the models which can be prohibitively expensive. We propose to improve the existing zero-reference low-light enhancement by leveraging the CLIP model to capture image prior and for semantic guidance. Specifically, we propose a data augmentation strategy to learn an image prior via prompt learning, based on image sampling, to learn the image prior without any need for paired or unpaired normal-light data. Next, we propose a semantic guidance strategy that maximally takes advantage of existing low-light annotation by introducing both content and context cues about the image training patches. We experimentally show, in a qualitative study, that the proposed prior and semantic guidance help to improve the overall image contrast and hue, as well as improve background-foreground discrimination, resulting in reduced over-saturation and noise over-amplification, common in related zero-reference methods. As we target machine cognition, rather than rely on assuming the correlation between human perception and downstream task performance, we conduct and present an ablation study and comparison with related zero-reference methods in terms of task-based performance across many low-light datasets, including image classification, object and face detection, showing the effectiveness of our proposed method.

低光增强CLIP图像处理语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。