arXiv:2412.07947cs.LGcs.AI2024-12被引 3

用向量符号架构视角,揭示GPT-2的层间计算机制

GPT-2 Through the Lens of Vector Symbolic Architectures

  • 将GPT-2的残差流视为向量符号运算,模拟正交捆绑与绑定
  • 发现模型权重中约70%可由向量符号原理解释
  • 适合对神经网络可解释性感兴趣的读者

理解Transformer模型的通用原理仍具挑战。通过稀疏自编码器(SAE)探测和解耦特征的实验表明,这些模型可能以方向形式在残差流中处理线性特征。本文探讨了仅解码器的Transformer架构与向量符号架构(VSA)之间的相似性,并通过实验表明,GPT-2采用近似正交的向量捆绑与绑定机制来实现层间计算与通信,类似VSA。研究进一步显示,这些原理能够解释模型中相当大一部分实际神经权重。

原文摘要 · Abstract (English)

Understanding the general priniciples behind transformer models remains a complex endeavor. Experiments with probing and disentangling features using sparse autoencoders (SAE) suggest that these models might manage linear features embedded as directions in the residual stream. This paper explores the resemblance between decoder-only transformer architecture and vector symbolic architectures (VSA) and presents experiments indicating that GPT-2 uses mechanisms involving nearly orthogonal vector bundling and binding operations similar to VSA for computation and communication between layers. It further shows that these principles help explain a significant portion of the actual neural weights.

Transformer可解释性向量符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。