提出空间感知的压缩框架,让视觉状态空间模型在降维时保持结构准确
Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

- 将降维重构为对空间单元的结构化操作,保留网格拓扑与邻域关系
- 在不训练的情况下,使VMamba的准确率提升63.3%,接近ViT性能
- 可直接接入现有流程,适用于多种视觉Mamba模型
Mamba在建模长视觉序列方面表现出色。然而,当对结构增强型Mamba变体应用标记压缩时,模型性能出现严重下降。我们归因于现有压缩方法的空间无感特性,违背了选择性扫描机制所需的二维结构前提。本文提出STORM——一种空间感知的标记压缩框架,旨在压缩过程中保持结构完整性。STORM将压缩重构为对空间单元的结构化操作,施加局部约束以维持网格拓扑和邻域一致性。作为即插即用模块,STORM在无需训练的情况下为现有压缩流程赋予显式空间感知能力。实验证明,STORM在多种视觉Mamba骨干网络下实现了训练自由设置下的最优剪枝精度。值得注意的是,其在VMamba上实现显著的准确率恢复,相较之前方法最高提升63.3%的top-1准确率;同时在PlainMamba上仅损失1.0%准确率,性能媲美ViT。
原文摘要 · Abstract (English)
Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a severe performance collapse. We attribute this degradation to the spatially agnostic nature of existing reduction methods, which violate the two-dimensional structural premise required by the selective scanning mechanism. In this work, we propose STORM, a spatial-aware token reduction framework designed to maintain structural integrity throughout the compression process. STORM reformulates reduction into a structured operation on spatial units, enforcing localized constraints to maintain both grid topology and neighborhood coherence. As a plug-and-play module, STORM equips existing reduction pipelines with explicit spatial awareness without any training. Empirical results demonstrate that STORM achieves state-of-the-art pruning accuracy across diverse vision Mamba backbones under training-free settings. Notably, STORM delivers a substantial accuracy recovery on VMamba, outperforming prior methods by up to 63.3\% in top-1 accuracy. Meanwhile, STORM incurs only a 1.0\% accuracy drop on PlainMamba, achieving performance comparable to ViT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。