用扩散模型引导状态空间模型,提升多模态显著性目标检测的边界精度。
DGSSM: Diffusion guided state-space models for multimodal salient object detection

- 将显著性检测建模为渐进去噪过程,融合扩散先验与多尺度状态空间编码。
- 在13个公开数据集上超越现有方法,边界精度显著提升,模型体积小。
- 适合需要高精度边界分割的多模态图像理解任务,如遥感、医疗图像分析。
显著性目标检测(SOD)需同时建模长程上下文依赖与细粒度结构细节,这对基于卷积、Transformer及基于Mamba的状态空间模型仍是挑战。尽管近期基于Mamba的方法实现高效全局推理,但往往难以恢复精确的目标边界。相比之下,扩散模型通过迭代去噪捕捉强结构先验,但其在判别性密集预测中的应用受限于计算成本与集成难题。本文提出DGSSM,一种扩散引导的状态空间(Mamba)框架,将多模态显著性检测建模为渐进去噪过程。该框架整合扩散结构先验与多尺度状态空间编码、自适应显著性提示及迭代Mamba去噪精修机制,以提升边界准确性。边界感知精修头与自蒸馏策略进一步增强空间连贯性与特征一致性。在13个公共基准测试(涵盖RGB、RGB-D、RGB-T设置)上的大量实验表明,DGSSM在多个评估指标上持续优于当前最优方法,同时保持紧凑模型规模。结果表明,扩散引导的状态空间建模是多模态密集预测任务中有效且通用的范式。
原文摘要 · Abstract (English)
Salient object detection (SOD) requires modeling both long-range contextual dependencies and fine-grained structural details, which remains challenging for convolutional, transformer-based, and Mamba-based state space models. While recent Mamba-based state space approaches enable efficient global reasoning, they often struggle to recover precise object boundaries. In contrast, diffusion models capture strong structural priors through iterative denoising, but their use in discriminative dense prediction is still limited due to computational cost and integration challenges. In this work, we propose DGSSM, a diffusion-guided state space (Mamba) framework that formulates multimodal salient object detection as a progressive denoising process. The framework integrates diffusion structural priors with multi-scale state space encoding, adaptive saliency prompting, and an iterative Mamba diffusion refinement mechanism to improve boundary accuracy. A boundary-aware refinement head and self-distillation strategy further enhance spatial coherence and feature consistency. Extensive experiments on 13 public benchmarks across RGB, RGB-D, and RGB-T settings demonstrate that DGSSM consistently outperforms state-of-the-art methods across multiple evaluation metrics while maintaining a compact model size. These results suggest that diffusion-guided state space modeling is an effective and generalizable paradigm for multimodal dense prediction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。