融合卷积与状态空间模型,提升遥感图像语义分割精度。
PPMamba: A Pyramid Pooling Local Auxiliary SSM-Based Model for Remote Sensing Image Semantic Segmentation
- 设计金字塔池化结构,多方向扫描特征图捕捉全局信息
- 在Vaihingen和LoveDA数据集上达到领先水平,优于现有模型
- 适合需要兼顾局部细节与长程依赖的遥感图像分析任务
语义分割是遥感领域的重要任务。然而,传统卷积神经网络(CNN)和基于Transformer的模型在捕捉长距离依赖关系方面存在局限,或计算成本过高。近期提出的先进状态空间模型(SSM),如Mamba,具有线性计算复杂度并能有效建模远距离依赖。尽管如此,基于Mamba的方法在保留局部语义信息方面仍面临挑战。为此,本文提出一种新网络——金字塔池化状态空间模型(PPMamba),将CNN与Mamba结合用于遥感图像语义分割。其核心模块为金字塔池化-状态空间模型(PP-SSM)块,结合局部辅助机制与全向状态空间模型(OSS),从八个方向选择性扫描特征图,全面捕获特征信息;同时,辅助分支采用金字塔形卷积结构,实现多尺度特征提取。在两个常用数据集ISPRS Vaihingen和LoveDA Urban上的大量实验表明,PPMamba性能媲美当前最优模型。
原文摘要 · Abstract (English)
Semantic segmentation is a vital task in the field of remote sensing (RS). However, conventional convolutional neural network (CNN) and transformer-based models face limitations in capturing long-range dependencies or are often computationally intensive. Recently, an advanced state space model (SSM), namely Mamba, was introduced, offering linear computational complexity while effectively establishing long-distance dependencies. Despite their advantages, Mamba-based methods encounter challenges in preserving local semantic information. To cope with these challenges, this paper proposes a novel network called Pyramid Pooling Mamba (PPMamba), which integrates CNN and Mamba for RS semantic segmentation tasks. The core structure of PPMamba, the Pyramid Pooling-State Space Model (PP-SSM) block, combines a local auxiliary mechanism with an omnidirectional state space model (OSS) that selectively scans feature maps from eight directions, capturing comprehensive feature information. Additionally, the auxiliary mechanism includes pyramid-shaped convolutional branches designed to extract features at multiple scales. Extensive experiments on two widely-used datasets, ISPRS Vaihingen and LoveDA Urban, demonstrate that PPMamba achieves competitive performance compared to state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。