arXiv:2511.18380cs.CV2025-11被引 2

揭示视觉Mamba模型的表征机制,证明其等价于低秩注意力。

RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models

  • 将Mamba视为Softmax注意力的低秩近似,统一了两种注意力形式。
  • 提出新量化指标,验证Mamba具备建模长距离依赖的能力。
  • 结合DINO预训练,激活图更清晰,适合可解释性研究。

Mamba近期作为视觉任务的有效骨干网络受到关注,但其在视觉领域的内在机制仍不明确。本文系统研究Mamba的表征特性,做出三项主要贡献:首先,理论分析表明Mamba可被视为Softmax注意力的低秩近似,从而弥合了Softmax与Linear注意力之间的表征差距;其次,提出一种新型二值分割评估指标,将激活图评估从定性拓展为定量,证实Mamba能有效建模长程依赖;第三,通过DINO进行自监督预训练,获得比传统监督方法更清晰的激活图,凸显Mamba在可解释性方面的潜力。值得注意的是,该模型在ImageNet上实现了78.5%的线性探测准确率,证明其强大性能。本工作希望为未来基于Mamba的视觉架构研究提供重要启示。

原文摘要 · Abstract (English)

Mamba has recently garnered attention as an effective backbone for vision tasks. However, its underlying mechanism in visual domains remains poorly understood. In this work, we systematically investigate Mamba's representational properties and make three primary contributions. First, we theoretically analyze Mamba's relationship to Softmax and Linear Attention, confirming that it can be viewed as a low-rank approximation of Softmax Attention and thereby bridging the representational gap between Softmax and Linear forms. Second, we introduce a novel binary segmentation metric for activation map evaluation, extending qualitative assessments to a quantitative measure that demonstrates Mamba's capacity to model long-range dependencies. Third, by leveraging DINO for self-supervised pretraining, we obtain clearer activation maps than those produced by standard supervised approaches, highlighting Mamba's potential for interpretability. Notably, our model also achieves a 78.5 percent linear probing accuracy on ImageNet, underscoring its strong performance. We hope this work can provide valuable insights for future investigations of Mamba-based vision architectures.

视觉Mamba注意力机制可解释性自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。