用CLIP模型提升水下图像增强的视觉真实感,避免过增强或欠增强。
Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement
- 引入CLIP感知损失,让增强图像更符合人眼视觉偏好。
- 通过课程对比正则化,动态优化不同难度负样本的约束效果。
- 在多个数据集上超越现有方法,尤其在泛化性和细节恢复上表现突出。
高质量水下图像对机器视觉任务和观赏性都至关重要。然而,光线吸收与散射严重降低了水下图像质量。基于深度学习的水下图像增强(UIE)方法虽取得良好效果,但常忽视人类感知,且解空间约束不足,导致增强后图像感知质量下降或内容失真。为此,本文提出一种结合对比语言-图像预训练(CLIP)感知损失模块与课程对比正则化的UIE方法。首先,利用CLIP模型的语义特征提取能力,学习适配水下图像的提示对,构建符合人类视觉感知的评价模型;该模型作为感知损失嵌入增强网络,提升输出图像的感知质量。其次,将该感知模型与课程对比正则化结合,强化在CLIP感知空间中对增强结果的约束,缓解欠增强与过增强风险。具体而言,通过CLIP模型评估并分类负样本的学习难度,实现对不同质量畸变图像的精细化利用。大量实验表明,本方法在视觉质量与泛化能力上均优于当前最优方法。
原文摘要 · Abstract (English)
High-quality underwater images are essential for both machine vision tasks and viewers with their aesthetic appeal.However, the quality of underwater images is severely affected by light absorption and scattering. Deep learning-based methods for Underwater Image Enhancement (UIE) have achieved good performance. However, these methods often overlook considering human perception and lack sufficient constraints within the solution space. Consequently, the enhanced images often suffer from diminished perceptual quality or poor content restoration.To address these issues, we propose a UIE method with a Contrastive Language-Image Pre-Training (CLIP) perception loss module and curriculum contrastive regularization. Above all, to develop a perception model for underwater images that more aligns with human visual perception, the visual semantic feature extraction capability of the CLIP model is leveraged to learn an appropriate prompt pair to map and evaluate the quality of underwater images. This CLIP perception model is then incorporated as a perception loss module into the enhancement network to improve the perceptual quality of enhanced images. Furthermore, the CLIP perception model is integrated with the curriculum contrastive regularization to enhance the constraints imposed on the enhanced images within the CLIP perceptual space, mitigating the risk of both under-enhancement and over-enhancement. Specifically, the CLIP perception model is employed to assess and categorize the learning difficulty level of negatives in the regularization process, ensuring comprehensive and nuanced utilization of distorted images and negatives with varied quality levels. Extensive experiments demonstrate that our method outperforms state-of-the-art methods in terms of visual quality and generalization ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。