arXiv:2410.00340cs.LGcs.AI2024-10被引 4

通过稀疏分解揭示GPT-2注意力头间通信机制

Sparse Attention Decomposition Applied to Circuit Tracing

  • 利用注意力头矩阵奇异向量的稀疏编码识别通信特征
  • 在IOI任务中追踪到冗余路径与具体通信特征
  • 为理解模型内部协作提供更精细的电路追踪方法

大量研究显示,注意力头协同完成复杂任务。通常认为注意力头间的通信依赖于将特定特征添加到标记残差中。本文旨在分离并识别GPT-2 small在间接宾语识别(IOI)任务中注意力头间用于协调的特征。关键突破在于发现这些特征常以注意力头矩阵奇异向量中的稀疏编码形式存在。我们分析了这些信号在各注意力头中的维度与出现频率。这种稀疏编码使信号能从残差背景中高效分离,并清晰识别注意力头间的通信路径。通过该方法追踪IOI任务的部分电路,揭示了此前研究未见的细节,阐明了GPT-2中存在的冗余路径特性,并首次识别出执行IOI时注意力头间传递的具体特征。

原文摘要 · Abstract (English)

Many papers have shown that attention heads work in conjunction with each other to perform complex tasks. It's frequently assumed that communication between attention heads is via the addition of specific features to token residuals. In this work we seek to isolate and identify the features used to effect communication and coordination among attention heads in GPT-2 small. Our key leverage on the problem is to show that these features are very often sparsely coded in the singular vectors of attention head matrices. We characterize the dimensionality and occurrence of these signals across the attention heads in GPT-2 small when used for the Indirect Object Identification (IOI) task. The sparse encoding of signals, as provided by attention head singular vectors, allows for efficient separation of signals from the residual background and straightforward identification of communication paths between attention heads. We explore the effectiveness of this approach by tracing portions of the circuits used in the IOI task. Our traces reveal considerable detail not present in previous studies, shedding light on the nature of redundant paths present in GPT-2. And our traces go beyond previous work by identifying features used to communicate between attention heads when performing IOI.

注意力机制模型可解释性神经网络追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。