arXiv:2601.10266cs.CL2026-01被引 4

用投影核度量注意力头权重子空间相似性,更清晰揭示模型内部结构。

Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel

  • 基于主角距离设计投影核,量化注意力头权重子空间的相似性。
  • 在IOI任务中,投影核比组合得分更准确反映头间交互关系。
  • 发现GPT2-small中第4层第7个头是核心枢纽,具有身份映射功能。

理解注意力头之间的关系对解析Transformer的内部结构至关重要,但现有度量方法未能有效捕捉这一结构。本文聚焦于由注意力头权重矩阵张成的子空间,利用基于主角的投影核(Projection Kernel, PK)来量化头与头之间的关系。实验表明,PK在IOI任务中比先前的组合得分等度量方法更清晰地再现了已知的头间交互。我们进一步提出一个框架,通过将PK分布与随机正交子空间生成的参考分布进行比较,以衡量其信息量。作为应用,我们基于PK构建有向图,发现GPT2-small中L4H7头充当枢纽,表现出身份映射功能。

原文摘要 · Abstract (English)

Understanding relationships between attention heads is essential for interpreting the internal structure of Transformers, yet existing metrics do not capture this structure well. We focus on the subspaces spanned by attention-head weight matrices and quantify head-to-head relationships using the Projection Kernel (PK), a principal-angle-based measure of subspace similarity. Experiments show that PK reproduces known head-to-head interactions on the IOI task more clearly than prior metrics such as the Composition Score. We further introduce a framework to quantify the informativeness of PK distributions by comparing them with a reference distribution derived from random orthogonal subspaces. As an application, we analyze a directed graph constructed from PK and show that, in GPT2-small, L4H7 acts as a hub by functioning as an identity head.

注意力机制子空间分析Transformer解释投影核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。