arXiv:2509.18683cs.CVcs.AI2025-09中稿 · ACM MM 2025被引 5

提出LEAF-Mamba模型,提升RGB-D显著物体检测的精度与效率。

LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection

  • 设计局部强调模块捕捉多尺度局部特征,增强细节感知。
  • 引入自适应融合模块实现双模态互补交互,提升融合效果。
  • 在16个主流方法中表现更优,且可迁移至红外检测任务。

RGB-D显著物体检测(SOD)旨在结合深度信息识别场景中最具突出性的物体。现有方法主要依赖受限于局部感受野的CNN,或计算复杂度呈二次增长的视觉变换器,难以兼顾性能与效率。最近的状态空间模型(SSM)如Mamba,以线性复杂度建模长程依赖展现出巨大潜力。然而直接将SSM应用于RGB-D SOD可能导致局部语义不足及跨模态融合不充分。为此,本文提出局部强调与自适应融合状态空间模型(LEAF-Mamba),包含两个新组件:1)局部强调状态空间模块(LE-SSM),用于捕获双模态的多尺度局部依赖;2)基于SSM的自适应融合模块(AFM),实现互补性跨模态交互与可靠集成。大量实验表明,LEAF-Mamba在有效性与效率上均持续优于16种前沿的RGB-D SOD方法。此外,该方法在RGB-T SOD任务上也取得优异表现,证明其强大的泛化能力。

原文摘要 · Abstract (English)

RGB-D salient object detection (SOD) aims to identify the most conspicuous objects in a scene with the incorporation of depth cues. Existing methods mainly rely on CNNs, limited by the local receptive fields, or Vision Transformers that suffer from the cost of quadratic complexity, posing a challenge in balancing performance and computational efficiency. Recently, state space models (SSM), Mamba, have shown great potential for modeling long-range dependency with linear complexity. However, directly applying SSM to RGB-D SOD may lead to deficient local semantics as well as the inadequate cross-modality fusion. To address these issues, we propose a Local Emphatic and Adaptive Fusion state space model (LEAF-Mamba) that contains two novel components: 1) a local emphatic state space module (LE-SSM) to capture multi-scale local dependencies for both modalities. 2) an SSM-based adaptive fusion module (AFM) for complementary cross-modality interaction and reliable cross-modality integration. Extensive experiments demonstrate that the LEAF-Mamba consistently outperforms 16 state-of-the-art RGB-D SOD methods in both efficacy and efficiency. Moreover, our method can achieve excellent performance on the RGB-T SOD task, proving a powerful generalization ability.

RGB-D检测状态空间模型跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。