arXiv:2505.24244cs.CLcs.LG2025-05ACL被引 4

用注意力剔除法解析Mamba模型的事实信息流动路径

Mamba Knockout for Unraveling Factual Information Flow

  • 借鉴Transformer的注意力剔除技术,分析Mamba模型的信息传递机制
  • 发现部分信息流动模式在不同模型间普遍存在,可能为大模型共性特征
  • 通过结构分解揭示特征对跨标记传递与单标记增强的不同作用

本文研究基于Mamba状态空间模型的语言模型中事实信息的流动规律。借助与基于Transformer架构及其注意力机制的理论与实证关联,我们将其原有的注意力可解释性技术——特别是注意力剔除法(Attention Knockout)——应用于Mamba-1和Mamba-2模型。通过该方法追踪信息在标记与层间的传播与定位,揭示了主语标记信息的涌现模式及逐层动态变化。值得注意的是,部分现象在不同模型间存在差异,而另一些则在所有被检视模型中普遍出现,暗示其可能为大语言模型的共性特征。进一步利用Mamba的结构化因子分解,我们解耦了不同‘特征’在实现标记间信息交换或增强单个标记表达方面的作用,从而提供了一个统一视角来理解Mamba内部运作机制。

原文摘要 · Abstract (English)

This paper investigates the flow of factual information in Mamba State-Space Model (SSM)-based language models. We rely on theoretical and empirical connections to Transformer-based architectures and their attention mechanisms. Exploiting this relationship, we adapt attentional interpretability techniques originally developed for Transformers--specifically, the Attention Knockout methodology--to both Mamba-1 and Mamba-2. Using them we trace how information is transmitted and localized across tokens and layers, revealing patterns of subject-token information emergence and layer-wise dynamics. Notably, some phenomena vary between mamba models and Transformer based models, while others appear universally across all models inspected--hinting that these may be inherent to LLMs in general. By further leveraging Mamba's structured factorization, we disentangle how distinct "features" either enable token-to-token information exchange or enrich individual tokens, thus offering a unified lens to understand Mamba internal operations.

模型解释Mamba信息流动可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。