arXiv:2607.01825cs.CV2026-07

针对水下图像模糊与色彩失真,提出物理引导的条件生成网络提升显著目标检测效果。

Rethinking Conditional Generation for Underwater Salient Object Detection

论文配图:Rethinking Conditional Generation for Underwater Salient Object Detection
图 1 · 摘自论文原文
  • 基于人类视觉设计多粒度模块,增强不同尺度模糊目标的检测能力。
  • 利用伪深度估计光衰减与后向散射,恢复色偏和边界模糊的图像特征。
  • 结合空间高斯先验与扩散变换器,有效抑制水下杂乱背景,适合复杂水下场景应用。

水下图像中的显著目标检测因低对比度、光照不均及散射吸收导致的颜色失真而面临挑战,限制了传统方法在水下环境的有效性。为此,我们提出退化感知条件生成网络(DCGNet),专为构建可靠的水下显著性生成条件特征而设计。首先,设计基于人类视觉系统的动态多粒度模块(DMG),以鲁棒检测具有模糊边界的各类尺度显著目标。其次,开发水下物理先验模块(UPP),利用伪深度引导估计水下光衰减与后向散射,恢复退化感知的RGB特征,缓解颜色失真与边界模糊。在此物理引导表征基础上,引入水下空间高斯模块(USG),从最强引导响应构建空间高斯显著性先验,增强以目标为中心的显著区域并抑制杂乱水下背景。此外,在去噪解码器中嵌入轻量级时间自适应扩散变压器(DiT)瓶颈,以在不同扩散步骤中精细化融合特征。在USOD10K、USOD、CSOD10K、MAS3K和RMAS上的全面实验表明,DCGNet显著优于现有最先进方法,验证其在复杂水下视觉任务中的潜力。

原文摘要 · Abstract (English)

Salient Object Detection in underwater images remains challenging due to low contrast, uneven illumination, and color distortion caused by scattering and absorption effects, which limit the effectiveness of conventional SOD methods in underwater environments. To address these challenges, we propose a Degradation-aware Conditional Generation Network (DCGNet), specifically designed to construct reliable conditional features for underwater saliency generation. First, we design a Dynamic Multi-Granularity module (DMG) grounded in the human visual system to robustly detect salient objects of varying scales with blurred boundaries. Then, we develop an Underwater Physics-Prior module (UPP), which utilizes pseudo-depth guidance to estimate underwater light attenuation and backscatter, thereby restoring degradation-aware RGB features and mitigating color distortion and boundary ambiguity. Based on the physics-guided representation, we introduce an Underwater Spatial Gaussian module (USG), which constructs a spatial Gaussian saliency prior from the strongest guided response to enhance object-centered salient regions and suppress cluttered underwater backgrounds. In addition, a lightweight timestep-adaptive Diffusion Transformer (DiT) bottleneck is inserted into the denoising decoder to refine fused features at different diffusion timesteps. Comprehensive experiments on USOD10K, USOD, CSOD10K, MAS3K, and RMAS demonstrate that DCGNet significantly outperforms existing state-of-the-art methods, verifying its potential for complex underwater visual applications.

水下检测条件生成物理先验扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。