arXiv:2602.22719cs.LG2026-02被引 2

发现并调控状态空间模型的激活瓶颈,提升性能且无需调参。

Interpreting and Steering State-Space Models via Activation Subspace Bottlenecks

  • 通过机制可解释性工具定位Mamba模型的激活子空间瓶颈。
  • 测试时仅缩放瓶颈激活,平均性能提升8.27%。
  • 提出Stable-Mamba架构,重训练后长上下文表现更优。

状态空间模型(SSMs)因其高效性成为构建语言模型的重要策略,避免了Transformer中注意力计算的二次复杂度。尽管前景广阔,现代SSMs的可解释性和可控性仍研究不足。本文通过机械可解释性工具,首次在Mamba系列模型中识别出激活子空间瓶颈,并提出一种测试时干预方法:仅对识别出的瓶颈激活进行标量乘法。在7个SSM模型和6个不同基准上,该方法平均提升性能8.27%,且无需任何任务特定调参。最后,通过修改瓶颈结构构建新架构Stable-Mamba,重新训练后在长上下文任务中实现显著性能提升。

原文摘要 · Abstract (English)

State-space models (SSMs) have emerged as an efficient strategy for building powerful language models, avoiding the quadratic complexity of computing attention in transformers. Despite their promise, the interpretability and steerability of modern SSMs remain relatively underexplored. We take a major step in this direction by identifying activation subspace bottlenecks in the Mamba family of SSM models using tools from mechanistic interpretability. We then introduce a test-time steering intervention that simply multiplies the activations of the identified bottlenecks by a scalar. Across 7 SSMs and 6 diverse benchmarks, this intervention improves performance by an average of 8.27%, without requiring any task-specific tuning. Finally, we validate that the identified bottlenecks are indeed hindering performance by modifying them to yield an architecture we call Stable-Mamba, which achieves long-context performance gains when retrained from scratch.

状态空间模型可解释性模型调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。