用范畴论解析注意力机制,揭示其数学本质。
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
- 将注意力的线性部分视为参数化1-态射,构建统一框架
- 多层注意力对应该函子的自由单子构造
- 适合研究注意力结构的数学与理论方向读者
自注意力机制革新了深度学习架构,但其核心数学结构仍不清晰。本文提出一个范畴论框架,聚焦自注意力的线性部分。我们证明查询、键、值映射自然构成2范畴$№{Para(Vect)}$中的参数化1-态射。在底层1范畴$№{Vect}$上,这些映射诱导出一个自函子,其迭代组合精确建模多层注意力。进一步证明,堆叠多个自注意力层等价于对该自函子构造自由单子。对于位置编码,我们表明严格加性嵌入对应仿射意义下的幺半群作用,而标准正弦编码虽非加性,但仍保有在单射(忠实)位置保持映射中的泛性质。还证明自注意力的线性部分对输入标记的置换具有自然等变性,并指出机制可解释性中识别出的“电路”可被解释为参数化1-态射的复合。这一范畴视角统一了几何、代数与可解释性分析方法,显式揭示了注意力的底层结构。全文仅限线性映射,非线性操作如softmax和层归一化因需更高级范畴构造而暂予保留。本工作扩展了近期深度学习范畴论基础研究,深化了对注意力代数结构的理解。
原文摘要 · Abstract (English)
Self-attention mechanisms have revolutionised deep learning architectures, yet their core mathematical structures remain incompletely understood. In this work, we develop a category-theoretic framework focusing on the linear components of self-attention. Specifically, we show that the query, key, and value maps naturally define a parametric 1-morphism in the 2-category $\mathbf{Para(Vect)}$. On the underlying 1-category $\mathbf{Vect}$, these maps induce an endofunctor whose iterated composition precisely models multi-layer attention. We further prove that stacking multiple self-attention layers corresponds to constructing the free monad on this endofunctor. For positional encodings, we demonstrate that strictly additive embeddings correspond to monoid actions in an affine sense, while standard sinusoidal encodings, though not additive, retain a universal property among injective (faithful) position-preserving maps. We also establish that the linear portions of self-attention exhibit natural equivariance to permutations of input tokens, and show how the "circuits" identified in mechanistic interpretability can be interpreted as compositions of parametric 1-morphisms. This categorical perspective unifies geometric, algebraic, and interpretability-based approaches to transformer analysis, making explicit the underlying structures of attention. We restrict to linear maps throughout, deferring the treatment of nonlinearities such as softmax and layer normalisation, which require more advanced categorical constructions. Our results build on and extend recent work on category-theoretic foundations for deep learning, offering deeper insights into the algebraic structure of attention mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。