提出自适应视觉状态空间模型,提升遥感图像显著目标检测效果
Beyond Global Scanning: Adaptive Visual State Space Modeling for Salient Object Detection in Optical Remote Sensing Images
- 用视觉状态空间编码器提取多尺度特征,增强长程依赖建模
- 在公开数据集上达到新最好结果,显著提升小目标与低对比度场景检测精度
- 适合遥感图像分析、地理信息识别等领域的研究人员参考
光学遥感图像中的显著目标检测(SOD)面临目标尺度变化大、前景与背景对比度低等挑战。现有基于视觉变换器(ViTs)和卷积神经网络(CNNs)的方法虽尝试融合全局与局部特征,但异构特征的有效整合仍受限。为此,本文提出自适应状态空间上下文网络(ASCNet),基于状态空间模型机制,同时捕捉长距离依赖并强化区域特征表示。具体地,采用视觉状态空间编码器提取多尺度特征;设计多层级上下文模块(MLCM),增强不同尺度特征间的跨层交互能力,提升模型结构感知力,更有效区分前景与背景;进一步设计自适应块状视觉状态空间(APVSS)作为解码器,融合动态自适应粒度扫描(DAGS)与粒度感知传播模块(GPM),对局部感知增强的特征图进行自适应分块扫描,充分获取局部区域信息,增强状态空间模型的局部建模能力。大量实验表明,该模型在多个公开遥感图像数据集上取得当前最优性能,验证了其有效性与优越性。
原文摘要 · Abstract (English)
Salient object detection (SOD) in optical remote sensing images (ORSIs) faces numerous challenges, including significant variations in target scales and low contrast between targets and the background. Existing methods based on vision transformers (ViTs) and convolutional neural networks (CNNs) architectures aim to leverage both global and local features, but the difficulty in effectively integrating these heterogeneous features limits their overall performance. To overcome these limitations, we propose an adaptive state space context network (ASCNet), which builds upon the state space model mechanism to simultaneously capture long-range dependencies and enhance regional feature representation. Specifically, we employ the visual state space encoder to extract multi-scale features. To further achieve deep guidance and enhancement of these features, we design a Multi-Level Context Module (MLCM), which module strengthens cross-layer interaction capabilities between features of different scales while enhancing the model's structural perception, allowing it to distinguish between foreground and background more effectively. Then, we design the Adaptive Patchwise Visual State Space (APVSS) block as the decoder of ASCNet, which integrates our proposed Dynamic Adaptive Granularity Scan (DAGS) and Granularity-aware Propagation Module (GPM). It performs adaptive patch scanning on feature maps enhanced by local perception, thereby capturing rich local region information and enhancing state space model's local modeling capability. Extensive experimental results demonstrate that the proposed model achieves state-of-the-art performance, validating its effectiveness and superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。