对比多种视觉状态空间模型,发现边界分割是主要短板
A Controlled Benchmark of Visual State-Space Backbones with Domain-Shift and Boundary Analysis for Remote-Sensing Segmentation

- 统一解码器和训练流程,只改变编码器进行公平比较
- 跨域泛化能力不对称,边界分割在分布偏移下表现差
- 提升效果更依赖鲁棒性设计而非单纯扩大模型规模
视觉状态空间模型(SSMs)被视作高效的视觉变换器替代方案,但其实际优势在现有研究中难以明确,因多数工作未将编码器影响与其他因素分离。本文构建了一个严格受控的基准测试,对比了代表性视觉SSM家族(VMamba、MambaVision、Spatial-Mamba)在遥感语义分割任务中的表现,仅编码器变化,其余部分保持一致。在LoveDA和ISPRS Potsdam数据集上,采用统一的四阶段特征接口和轻量级解码器进行评估,结果揭示三个核心发现:同族模型规模扩展仅带来微弱收益;跨域泛化能力显著不对称;分布偏移下边界分割是主要失败模式。尽管视觉SSMs相较所考察的控制型CNN与Transformer基线表现出更优的精度-效率权衡,但结果表明未来改进应更聚焦于鲁棒性设计与边界感知解码,而非仅依赖编码器规模扩展。本研究通过统一可复现协议隔离编码器行为,为后续基于Mamba的分割主干设计与评估提供了实用基准。
原文摘要 · Abstract (English)
Visual state-space models (SSMs) are increasingly promoted as efficient alternatives to Vision Transformers, yet their practical advantages remain unclear under fair comparison because existing studies rarely isolate encoder effects from decoder and training choices. We present a strictly controlled benchmark of representative visual SSM families, including VMamba, MambaVision, and Spatial-Mamba, for remote-sensing semantic segmentation, in which only the encoder varies across experiments. Evaluated on LoveDA and ISPRS Potsdam under a unified 4-stage feature interface and a fixed lightweight decoder, the benchmark reveals three main findings, intra-family scaling yields only modest gains, cross-domain generalization is strongly asymmetric, and boundary delineation is the dominant failure mode under distribution shift. Although visual SSMs achieve favorable accuracy-efficiency trade-offs relative to the controlled CNN and Transformer baselines considered here, the results suggest that future improvements are more likely to come from robustness-oriented design and boundary-aware decoding than from encoder scaling alone. By isolating encoder behavior under a unified and reproducible protocol, this study establishes a practical reference benchmark for the design and evaluation of future Mamba-based segmentation backbones
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。