用物理模型与文本语义联合增强水下图像,效果优于主流方法。
Retinex Meets Language: A Physics-Semantics-Guided Underwater Image Enhancement Network
- 结合Retinex光照校正与CLIP生成的文本描述进行图像修复
- 在6418对图文数据上训练,提升颜色还原与视觉感知质量
- 首次引入文本引导和多模态数据集,适合水下视觉研究者
水下图像常因光吸收与散射导致颜色失真、对比度低、可见度差。现有水下图像增强方法分为基于先验和基于学习两类:前者依赖刚性物理假设,适应性差;后者受限于数据稀缺与泛化能力弱。为此,本文提出一种物理-语义协同的水下图像增强网络(PSG-UIENet),融合基于Retinex的无先验光照估计器与语义引导的图像恢复器。恢复器利用对比语言图像预训练(CLIP)模型生成的文本描述,注入高层语义信息以实现感知有意义的引导。由于缺乏公开的多模态水下图像数据集,本文构建了大规模图像-文本配对数据集LUIQD-TD,包含6,418组图像-参考文本三元组。为显式衡量并优化图文语义一致性,设计了图像-文本语义相似性(ITSS)损失函数。实验证明,该方法在自建数据集及四个公开数据集上均超越或媲美15种先进方法。
原文摘要 · Abstract (English)
Underwater images often suffer from severe degradation caused by light absorption and scattering, leading to color distortion, low contrast and reduced visibility. Existing Underwater Image Enhancement (UIE) methods can be divided into two categories, i.e., prior-based and learning-based methods. The former rely on rigid physical assumptions that limit the adaptability, while the latter often face data scarcity and weak generalization. To address these issues, we propose a Physics-Semantics-Guided Underwater Image Enhancement Network (PSG-UIENet), which couples the Retinex-grounded illumination correction with the language-informed guidance. This network comprises a Prior-Free Illumination Estimator and a Semantics-Guided Image Restorer. In particular, the restorer leverages the textual descriptions generated by the Contrastive Language-Image Pre-training (CLIP) model to inject high-level semantics for perceptually meaningful guidance. Since multimodal UIE data sets are not publicly available, we also construct a large-scale image-text UIE data set, namely, LUIQD-TD, which contains 6,418 image-reference-text triplets. To explicitly measure and optimize semantic consistency between textual descriptions and images, we further design an Image-Text Semantic Similarity (ITSS) loss function. To our knowledge, this study makes the first effort to introduce both textual guidance and the multimodal data set into UIE tasks. Extensive experiments on our data set and four publicly available data sets demonstrate that the proposed PSG-UIENet achieves superior or comparable performance against fifteen state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。