arXiv:2606.07604cs.LGcs.AI2026-06

提出贡献权重,更准确衡量注意力中词元的重要程度

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

论文配图:Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
图 1 · 摘自论文原文
  • 基于投影的度量方法,融合注意力权重、值向量大小和方向对齐
  • 在多个模型和任务中,比传统注意力权重更准识别关键词元
  • 揭示注意力黑洞主动抑制信息,稳定低置信度词元表示

分析注意力权重已成为解释大语言模型信息流的标准方法,但该方法忽视了被聚合的值向量的几何特性。为此,我们引入了 extit{贡献权重},一种基于投影的度量,通过考虑词元的注意力权重、值向量幅值及其与层输出的方向对齐,量化其影响力。实验表明,贡献权重能更真实地反映词元重要性,在不同解码器模型、任务和数据集上均优于基于注意力的度量。此外,该度量揭示了 extit{注意力黑洞}的主动功能:此前认为是被动存储过量注意力,而实际上通过凹凸关系抑制信息,对抗低置信度词元的语义漂移,稳定表示。

原文摘要 · Abstract (English)

Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approach has significant limitations as it neglects the geometric properties of the value vectors being aggregated. To address this gap, we introduce \emph{Contribution Weights}, a projection-based metric that quantifies a token's influence by accounting for it's attention weight, value magnitude, and directional alignment with the layer output. We demonstrate that contribution weights provide a more faithful measure of token importance, consistently outperforming attention-based metrics in identifying semantically critical tokens across different decoder-only models, tasks, and datasets. Further, our metric enables novel mechanistic analysis of \emph{attention sinks}. While previous work characterized sinks as passive repositories for excess attention, we reveal they serve an active functional role, suppressing information through a convex relationship between sink rate and output norm, stabilizing representations by opposing the semantic drift of low-confidence tokens.

注意力机制可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。