用视觉状态空间模型提升遥感图像分割精度与效率
Remote Sensing Image Segmentation Using Vision Mamba and Multi-Scale Multi-Frequency Feature Fusion
- 融合跨2D扫描与卷积分支,兼顾全局与局部特征提取
- 通过多频多尺度融合模块提升细节信息利用与特征表达能力
- 在多个遥感数据集上超越主流算法,兼具高精度与低计算开销
随着遥感成像技术的发展,如何高效处理高分辨率、多样化的卫星图像以提升分割精度和解读效率成为关键研究方向。尽管基于CNN和Transformer的分割算法已取得显著进展,但在分割精度与计算复杂度之间仍难以平衡,制约了其实际应用。为此,本文提出一种基于视觉状态空间模型(SSM)的新型混合语义分割网络——CVMH-UNet。该方法设计了交叉2D扫描视觉状态空间模块(CVSSBlock),通过跨方向2D扫描全面捕获全局信息;同时引入卷积分支,弥补视觉马尔可夫模型(VMamba)在局部特征获取上的不足,实现对全局与局部特征的协同分析。此外,为解决直接跳跃连接带来的区分力弱与融合不充分问题,提出了多频多尺度特征融合模块(MFMSBlock)。该模块利用二维离散余弦变换(2D DCT)引入多频信息,通过逐点卷积分支提供额外尺度的局部细节信息,并沿通道维度聚合多尺度特征,实现精细化融合。在多个知名遥感图像数据集上的实验表明,所提方法在保持低计算复杂度的同时,实现了更优的分割性能,优于当前先进算法。
原文摘要 · Abstract (English)
As remote sensing imaging technology continues to advance and evolve, processing high-resolution and diversified satellite imagery to improve segmentation accuracy and enhance interpretation efficiency emerg as a pivotal area of investigation within the realm of remote sensing. Although segmentation algorithms based on CNNs and Transformers achieve significant progress in performance, balancing segmentation accuracy and computational complexity remains challenging, limiting their wide application in practical tasks. To address this, this paper introduces state space model (SSM) and proposes a novel hybrid semantic segmentation network based on vision Mamba (CVMH-UNet). This method designs a cross-scanning visual state space block (CVSSBlock) that uses cross 2D scanning (CS2D) to fully capture global information from multiple directions, while by incorporating convolutional neural network branches to overcome the constraints of Vision Mamba (VMamba) in acquiring local information, this approach facilitates a comprehensive analysis of both global and local features. Furthermore, to address the issue of limited discriminative power and the difficulty in achieving detailed fusion with direct skip connections, a multi-frequency multi-scale feature fusion block (MFMSBlock) is designed. This module introduces multi-frequency information through 2D discrete cosine transform (2D DCT) to enhance information utilization and provides additional scale local detail information through point-wise convolution branches. Finally, it aggregates multi-scale information along the channel dimension, achieving refined feature fusion. Findings from experiments conducted on renowned datasets of remote sensing imagery demonstrate that proposed CVMH-UNet achieves superior segmentation performance while maintaining low computational complexity, outperforming surpassing current leading-edge segmentation algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。