针对水下图像结构线索可靠性差异,提出自适应协调机制提升显著目标检测精度。
Learning Spatially Adaptive Structural Coordination for Underwater Salient Object Detection

- 构建边界敏感与区域一致两类互补结构表征
- 在USOD10K和USOD上分别降低MAE 4.07%和23.53%
- 轻量模型可在Jetson TX2 NX上达21 FPS,适合水下机器人部署
水下显著目标检测(USOD)在水下场景理解与视觉引导机器人应用中日益重要。然而,水下图像的空间非均匀退化导致结构线索的可靠性随位置变化:边界敏感响应虽能增强目标轮廓,但易受退化噪声影响;区域一致响应虽提升语义完整性,却可能模糊边界。现有方法极少显式考虑水下图像退化下的结构线索可靠性空间差异。为此,本文提出SASC-USOD框架,通过学习空间自适应结构协调实现USOD。该框架构建两种互补结构表示:边界敏感表示通过固定拉普拉斯滤波与可学习局部细节变换结合,增强判别性边界信息;区域一致表示则通过双范围各向异性大核上下文聚合,捕捉长程结构一致性。随后引入空间协调模块,估计两类表示的相对可靠性,并根据图像内容自适应协调其贡献。在USOD10K与USOD基准上的大量实验表明,SASC-USOD持续优于现有方法,相较最强竞争方法分别降低MAE 4.07%和23.53%。此外,其轻量变体在NVIDIA Jetson TX2 NX上运行达21 FPS,具备在轨水下机器人感知能力。
原文摘要 · Abstract (English)
Underwater salient object detection (USOD) has attracted increasing attention for underwater scene understanding and vision-guided robotic applications. However, the spatially non-uniform degradation in underwater images causes spatially varying reliability of structural cues: boundary-sensitive responses can enhance object contours but are vulnerable to degradation-induced noise, whereas region-coherent responses improve semantic completeness but may blur object boundaries. Existing methods rarely explicitly consider the spatial variation in structural cue reliability under underwater image degradation. To address this problem, this work proposes SASC-USOD, a novel framework for learning spatially adaptive structural coordination in USOD. The proposed framework constructs two complementary structural representations with different characteristics. A boundary-sensitive representation is obtained by combining fixed Laplacian filtering with a learnable local-detail transformation to enhance discriminative boundary information, while a region-coherent representation is generated through dual-range anisotropic large-kernel contextual aggregation to capture long-range structural consistency. A spatial coordination module is then introduced to estimate the relative reliability of these structural representations and adaptively coordinate their contributions according to image content. Extensive experiments on the USOD10K and USOD benchmarks demonstrate that SASC-USOD consistently outperforms existing methods, reducing MAE by 4.07\% and 23.53\% compared with the strongest competing method, respectively. Moreover, its lightweight variant runs at 21 FPS on an NVIDIA Jetson TX2 NX, demonstrating its capability for onboard underwater robotic perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。