arXiv:2409.16604cs.CV2024-09被引 8

用半监督学习提升暗光图像增强,利用未配对数据训练更自然的光照效果。

Semi-LLIE: Semi-supervised Contrastive Learning with Mamba-based Low-light Image Enhancement

  • 基于均值教师框架,融合未配对数据进行训练。
  • 引入语义感知对比损失,还原真实光照分布,避免颜色偏移。
  • 结合Mamba架构与RAM感知损失,生成纹理更丰富的增强图像。

尽管低光图像增强技术取得了显著进展,但成对数据的稀缺已成为进一步发展的主要障碍。本文提出一种基于均值教师的半监督低光图像增强(Semi-LLIE)框架,将未配对数据融入模型训练。传统均值教师方法在该任务中面临两大挑战:一是像素级一致性损失难以有效传递真实的光照分布,导致增强图像出现色偏;二是先进增强方法难以与均值教师框架协同,忽略局部结构信息建模,影响暗区细节恢复。为此,我们提出语义感知对比损失,以准确传递光照分布,实现自然色彩增强;设计基于Mamba的骨干网络,结合多尺度特征学习,强化局部像素关系建模能力,提升纹理细节;并提出基于大规模视觉-语言模型RAM的感知损失,进一步丰富生成图像的文本细节。实验表明,Semi-LLIE在定量和定性指标上均优于现有方法。

原文摘要 · Abstract (English)

Despite the impressive advancements made in recent low-light image enhancement techniques, the scarcity of paired data has emerged as a significant obstacle to further advancements. This work proposes a mean-teacher-based semi-supervised low-light enhancement (Semi-LLIE) framework that integrates the unpaired data into model training. The mean-teacher technique is a prominent semi-supervised learning method, successfully adopted for addressing high-level and low-level vision tasks. However, two primary issues hinder the naive mean-teacher method from attaining optimal performance in low-light image enhancement. Firstly, pixel-wise consistency loss is insufficient for transferring realistic illumination distribution from the teacher to the student model, which results in color cast in the enhanced images. Secondly, cutting-edge image enhancement approaches fail to effectively cooperate with the mean-teacher framework to restore detailed information in dark areas due to their tendency to overlook modeling structured information within local regions. To mitigate the above issues, we first introduce a semantic-aware contrastive loss to faithfully transfer the illumination distribution, contributing to enhancing images with natural colors. Then, we design a Mamba-based low-light image enhancement backbone to effectively enhance Mamba's local region pixel relationship representation ability with a multi-scale feature learning scheme, facilitating the generation of images with rich textural details. Further, we propose novel perceptive loss based on the large-scale vision-language Recognize Anything Model (RAM) to help generate enhanced images with richer textual details. The experimental results indicate that our Semi-LLIE surpasses existing methods in both quantitative and qualitative metrics.

低光增强半监督Mamba感知损失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。