arXiv:2505.19895cs.CV2025-05被引 4

用视觉语言模型提升水下图像增强,更真实且细节丰富。

Underwater Diffusion Attention Network with Contrastive Language-Image Joint Learning for Underwater Image Enhancement

  • 结合视觉语言模型与空间注意力,精准修复水下雾霾和色彩偏移。
  • 在多个水下数据集上优于现有方法,保持语义一致性并还原自然外观。
  • 适合需要高保真水下图像的科研与海洋探测应用。

水下图像常受光吸收、散射、色偏和伪影等复杂退化影响,严重影响水下目标检测、识别与场景理解。现有基于扩散模型的方法多依赖合成配对数据集,因真实水下参考数据稀缺,导致模型引入偏差且泛化能力受限。此外,微调过程易破坏已学习先验,引发不真实增强。为此,我们提出UDAN-CLIP,一种基于合成水下数据预训练的图像到图像扩散框架,融合定制分类器(基于视觉语言模型)、空间注意力模块及新型CLIP-Diffusion损失。分类器保留自然陆地图像先验并语义引导扩散过程;空间注意力模块聚焦局部退化如雾气与低对比度;CLIP-Diffusion损失强化视觉-文本对齐,保障增强过程中的语义一致性。实验表明,该模型在定量指标与定性视觉对比中均表现优异,有效纠正退化并恢复自然外观,尤其在挑战性水下条件下表现突出。

原文摘要 · Abstract (English)

Underwater images are often affected by complex degradations such as light absorption, scattering, color casts, and artifacts, making enhancement critical for effective object detection, recognition, and scene understanding in aquatic environments. Existing methods, especially diffusion-based approaches, typically rely on synthetic paired datasets due to the scarcity of real underwater references, introducing bias and limiting generalization. Furthermore, fine-tuning these models can degrade learned priors, resulting in unrealistic enhancements due to domain shifts. To address these challenges, we propose UDAN-CLIP, an image-to-image diffusion framework pre-trained on synthetic underwater datasets and enhanced with a customized classifier based on vision-language model, a spatial attention module, and a novel CLIP-Diffusion loss. The classifier preserves natural in-air priors and semantically guides the diffusion process, while the spatial attention module focuses on correcting localized degradations such as haze and low contrast. The proposed CLIP-Diffusion loss further strengthens visual-textual alignment and helps maintain semantic consistency during enhancement. The proposed contributions empower our UDAN-CLIP model to perform more effective underwater image enhancement, producing results that are not only visually compelling but also more realistic and detail-preserving. These improvements are consistently validated through both quantitative metrics and qualitative visual comparisons, demonstrating the model's ability to correct distortions and restore natural appearance in challenging underwater conditions.

水下图像扩散模型视觉语言图像增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。