arXiv:2503.02270cs.CV2025-03被引 1

用状态空间模型融合多模态信息,提升低质量深度图下的显著性目标检测效果。

SSNet: Saliency Prior and State Space Model-based Network for Salient Object Detection in RGB-D Images

  • 引入状态空间模型捕捉跨模态全局依赖,计算效率高。
  • 在7个基准数据集上均超越现有方法,性能领先。
  • 适合处理深度图质量差的场景,如机器人视觉与增强现实。

RGB-D图像中的显著性目标检测(SOD)是计算机视觉的关键任务,广泛应用于场景理解、机器人和增强现实。然而,现有方法难以捕捉模态间的全局依赖关系,缺乏来自RGB与深度数据的完整显著性先验,且对低质量深度图表现不佳。为此,本文提出SSNet,一种基于显著性先验与状态空间模型(SSM)的RGB-D SOD网络。不同于传统卷积或Transformer结构,SSNet采用线性复杂度的多模态多尺度解码器模块,通过跨模态选择性扫描状态空间模型(CM-S6)有效建模模态间全局依赖。同时,设计显著性增强模块(SEM),融合三种显著性先验与深层特征,提升显著对象定位精度。针对低质量深度图问题,提出自适应对比度增强技术,动态优化深度图以适配任务需求。在7个基准数据集上的大量定量与定性实验表明,SSNet显著优于当前最优方法。

原文摘要 · Abstract (English)

Salient object detection (SOD) in RGB-D images is an essential task in computer vision, enabling applications in scene understanding, robotics, and augmented reality. However, existing methods struggle to capture global dependency across modalities, lack comprehensive saliency priors from both RGB and depth data, and are ineffective in handling low-quality depth maps. To address these challenges, we propose SSNet, a saliency-prior and state space model (SSM)-based network for the RGB-D SOD task. Unlike existing convolution- or transformer-based approaches, SSNet introduces an SSM-based multi-modal multi-scale decoder module to efficiently capture both intra- and inter-modal global dependency with linear complexity. Specifically, we propose a cross-modal selective scan SSM (CM-S6) mechanism, which effectively captures global dependency between different modalities. Furthermore, we introduce a saliency enhancement module (SEM) that integrates three saliency priors with deep features to refine feature representation and improve the localization of salient objects. To further address the issue of low-quality depth maps, we propose an adaptive contrast enhancement technique that dynamically refines depth maps, making them more suitable for the RGB-D SOD task. Extensive quantitative and qualitative experiments on seven benchmark datasets demonstrate that SSNet outperforms state-of-the-art methods.

显著性检测多模态状态空间模型深度图增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。