用视觉语言模型让水下图像增强更关注关键物体,提升下游任务效果。
Empowering Semantic-Sensitive Underwater Image Enhancement with VLM

- 通过VLM生成物体描述,反向生成语义引导图。
- 双机制引导增强网络,聚焦关键区域修复。
- 显著提升检测与分割性能,适合实际应用。
近年来,基于学习的水下图像增强(UIE)技术快速发展。然而,高质量增强结果与自然图像之间的分布差异会阻碍下游视觉任务中的语义线索提取,限制现有增强模型的适应性。为此,本文提出一种新学习机制,利用视觉语言模型(VLM)赋予UIE模型语义敏感能力。具体而言,先通过VLM对退化图像中的关键物体生成文本描述,再使用文本-图像对齐模型将这些描述映射回图像,生成空间语义引导图。该图通过交叉注意力与显式对齐损失双重机制,引导UIE网络在重建时聚焦于语义敏感区域,而非追求全局均匀改善,从而确保关键物体特征的忠实还原。实验表明,将该策略应用于不同UIE基线模型后,显著提升了感知质量指标,并增强了目标检测与分割任务的表现,验证了其有效性与适应性。
原文摘要 · Abstract (English)
In recent years, learning-based underwater image enhancement (UIE) techniques have rapidly evolved. However, distribution shifts between high-quality enhanced outputs and natural images can hinder semantic cue extraction for downstream vision tasks, thereby limiting the adaptability of existing enhancement models. To address this challenge, this work proposes a new learning mechanism that leverages Vision-Language Models (VLMs) to empower UIE models with semantic-sensitive capabilities. To be concrete, our strategy first generates textual descriptions of key objects from a degraded image via VLMs. Subsequently, a text-image alignment model remaps these relevant descriptions back onto the image to produce a spatial semantic guidance map. This map then steers the UIE network through a dual-guidance mechanism, which combines cross-attention and an explicit alignment loss. This forces the network to focus its restorative power on semantic-sensitive regions during image reconstruction, rather than pursuing a globally uniform improvement, thereby ensuring the faithful restoration of key object features. Experiments confirm that when our strategy is applied to different UIE baselines, significantly boosts their performance on perceptual quality metrics as well as enhances their performance on detection and segmentation tasks, validating its effectiveness and adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。