用注意力传播分析Transformer内部特征演化,揭示深度增强的注意力收缩现象。
Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

- 引入内部可解释性,通过注意力滚动追踪特征在层间的传播路径。
- 深度越深,注意力传播收缩越强,表明高层更聚焦关键特征。
- 适合关注模型内部机制的开发者与研究者,尤其关注注意力行为分析。
特征归因方法为输入变量与模型输出之间的关联分配得分,但无法直接刻画显式定义的交互算子在中间层如何组合。本文提出‘内部可解释性’,从传播视角理解模型内部结构,并以表格型Transformer的注意力滚动为例进行实例化。将滚动视为行随机算子,编码特征标记间的注意力驱动传播。基于Doeblin-Dobrushin收缩理论,证明具有小多布林系数的滚动算子在数值上接近秩一随机矩阵,其公共行由归一化列和决定。该结果赋予滚动传播轮廓结构性解释。在代谢组学年龄预测任务中,测量到的滚动收缩随深度增强。训练后与随机初始化模型表现出不同传播轮廓,但当前实验未确立单个滚动排序变量的预测相关性。与PCA及GradientExplainer对SHAP的近似比较显示,高排序变量间存在局部一致性,但完整排序整体一致性弱。因此,注意力滚动在此作为注意力传播的诊断工具,而非因果解释或完整Transformer的忠实归因。
原文摘要 · Abstract (English)
Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretability}, a propagation-based perspective on internal model organization, and instantiate it for tabular Transformers using attention rollout. We interpret rollout as a row-stochastic operator encoding attention-mediated propagation between feature tokens. By applying classical Doeblin--Dobrushin contraction theory, we show that a rollout operator with a small Dobrushin coefficient is quantitatively close to a rank-one stochastic matrix whose common row is determined by its normalized column sums. This result gives a structural interpretation to the corresponding rollout propagation profile. In Transformers trained for metabolomic age prediction, the measured rollout contraction strengthens with depth. Trained and randomly initialized models also exhibit different propagation profiles, although the present experiments do not establish the predictive relevance of individual rollout-ranked variables. Exploratory comparisons with PCA and GradientExplainer approximations to SHAP reveal localized agreement among highly ranked variables but weak agreement across complete rankings. Attention rollout is therefore used here as a diagnostic of attention-mediated propagation, not as a causal explanation or faithful attribution of the complete Transformer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。