arXiv:2506.15541math.NAcs.AI2025-06

揭示Transformer注意力机制的内在与外在结构,为模型可解释性和剪枝提供理论基础。

Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity

  • 通过微分几何分析证明注意力机制对Softmax不变,依赖头内结构组织
  • 构建查询/键/头三重张量的层次树,发现网络具有可利用的规律性结构
  • 基于哈尔基函数展开实现网络稀疏量化,适用于剪枝与架构对比

本文研究Transformer自注意力机制的内在(注意力头内部)与外在(注意力头之间)结构。通过拟微分计算理论证明自注意力机制对Softmax激活具有不变性,该性质依赖于注意力头的内在组织结构。进一步采用现有张量层次组织方法,构建以查询、键和头轴为基础的层级划分树,揭示网络三阶张量在几何空间中的规律性特征。这一组织结构使常见信号处理任务得以有效执行。通过可视化注意力头的层次树与扩散映射嵌入进行定性展示,并定量分析各注意力头及全网在查询、键、头空间上分别基于双哈尔与三哈尔基函数的展开系数,评估网络稀疏性。通过视觉与语言模型的计算实例验证理论与方法的有效性。研究成果带来双重启示:(1)为后续可解释性分析提供理论依据,可实用于下游任务;(2)可利用三阶张量组织实现模型剪枝(基于稀疏性)与架构比较。

原文摘要 · Abstract (English)

We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for invariance of the self-attention mechanism to softmax activation is obtained by appealing to paradifferential calculus, (and is supported by computational examples), which relies on the intrinsic organization of the attention heads. Furthermore, we use an existing methodology for hierarchical organization of tensors to examine network structure by constructing hierarchal partition trees with respect to the query, key, and head axes of network 3-tensors. Such an organization is consequential since it allows one to profitably execute common signal processing tasks on a geometry where the organized network 3-tensors exhibit regularity. We exemplify this qualitatively, by visualizing the hierarchical organization of the tree comprised of attention heads and the diffusion map embeddings, and quantitatively by investigating network sparsity with the expansion coefficients of individual attention heads and the entire network with respect to the bi and tri-haar bases (respectively) on the space of queries, keys, and heads of the network. To showcase the utility of our theoretical and methodological findings, we provide computational examples using vision and language transformers. The ramifications of these findings are two-fold: (1) a subsequent step in interpretability analysis is theoretically admitted, and can be exploited empirically for downstream interpretability tasks (2) one can use the network 3-tensor organization for empirical network applications such as model pruning (by virtue of network sparsity) and network architecture comparison.

Transformer注意力机制可解释性稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。