揭示注意力黑洞是变压器的几何参考系原理,非缺陷而是稳定坐标系的必然结果。
What are you sinking? A geometric approach on attention sink
- 从几何视角解释注意力黑洞,源于建立高维空间参考系的需要。
- 识别出三种参考系类型:集中式、分布式、双向式,训练初期即出现。
- 位置编码等组件决定参考系类型,为模型设计提供新思路。
注意力黑洞(AS)是变压器注意力图中的稳定模式,特定标记(常为特殊标记或位置锚点)会吸引大量其他标记的注意力。我们证明,在变压器中,AS并非架构缺陷,而是基本几何原理的表现:建立锚定表征空间的参考系。分析多种架构后,识别出三种与注意力黑洞相关的参考系类型——集中式、分布式和双向式,它们在训练初期作为高维空间中建立稳定坐标系统的最优解出现。研究还揭示了架构组件(特别是位置编码实现方式)对参考系类型的影响。这一视角重构了对变压器注意力机制的理解,并为架构设计及注意力黑洞关系提供了新见解。
原文摘要 · Abstract (English)
Attention sink (AS) is a consistent pattern in transformer attention maps where certain tokens (often special tokens or positional anchors) disproportionately attract attention from other tokens. We show that in transformers, AS is not an architectural artifact, but it is the manifestation of a fundamental geometric principle: the establishment of reference frames that anchor representational spaces. We analyze several architectures and identify three distinct reference frame types, centralized, distributed, and bidirectional, that correlate with the attention sink phenomenon. We show that they emerge during the earliest stages of training as optimal solutions to the problem of establishing stable coordinate systems in high-dimensional spaces. We show the influence of architecture components, particularly position encoding implementations, on the specific type of reference frame. This perspective transforms our understanding of transformer attention mechanisms and provides insights for both architecture design and the relationship with AS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。