arXiv:2508.13599cs.CV2025-08

提出新方法提升视觉状态空间模型效率,减少计算量同时保持性能。

Towards Efficient Vision State Space Models via Token Merging

  • 用状态转移参数衡量令牌重要性,动态合并低信息量令牌。
  • 在极端压缩下仍保持性能稳定,优于现有方法。
  • 适用于图像、视频和音频多模态任务,通用性强。

状态空间模型(SSMs)在计算机视觉中展现出强大能力,但提升其计算效率对实际部署至关重要。虽然令牌缩减可有效提升效率,但应用于SSM时需兼顾其独特的序列建模特性。本文提出针对基于SSM的视觉模型的令牌合并策略MaMe,解决两个关键问题:量化令牌重要性与保留序列属性。该方法利用状态转移参数Δ作为信息量度量,并引入策略性令牌排列以保持序列信息流动。大量实验表明,MaMe在微调和现成模型上均实现更优的效率-性能权衡。尤其在激进令牌压缩下,现有方法性能显著下降,而本方法仍保持鲁棒性。此外,MaMe在图像分类之外,展现出跨视频与音频领域的强泛化能力,为多种SSM应用提供了高效的扩展方案。

原文摘要 · Abstract (English)

State Space Models (SSMs) have emerged as powerful architectures in computer vision, yet improving their computational efficiency remains crucial for practical and scalable deployment.While token reduction serves as an effective approach for model efficiency, applying it to SSMs requires careful consideration of their unique sequential modeling capabilities.In this work, we propose MaMe, a token-merging strategy tailored for SSM-based vision models.MaMe addresses two key challenges: quantifying token importance and preserving sequential properties. Our approach leverages the state transition parameter $\mathbfΔ$ as an informativeness measure and introduces strategic token arrangements to preserve sequential information flow.Extensive experiments demonstrate that MaMe achieves superior efficiency-performance trade-offs for both fine-tuned and off-the-shelf models. Particularly, our approach maintains robustness even under aggressive token reduction where existing methods undergo significant performance degradation.Beyond image classification, MaMe shows strong generalization capabilities across video and audio domains, establishing an effective approach for enhancing efficiency in diverse SSM applications.

视觉模型状态空间效率优化令牌合并

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。