提出可解析的注意力模型,揭示深度注意力层的学习规律。
Bayes optimal learning of attention-indexed models
- 基于双线性交互构建可解析的注意力理论框架
- 首次给出贝叶斯最优泛化误差的闭式解并发现相变现象
- 适合研究自注意力机制的理论学者与模型设计者
我们提出注意力索引模型(AIM),一种用于分析深度注意力层学习的理论框架。受多索引模型启发,AIM 揭示了标记级输出如何从高维嵌入的分层双线性交互中产生。与以往可解析的注意力模型不同,AIM 允许全宽的键和查询矩阵,更贴近实际 Transformer 架构。利用统计物理和随机矩阵论工具,我们推导出贝叶斯最优泛化误差的闭式表达,并揭示了样本复杂度、模型宽度和序列长度之间的尖锐相变。我们提出一种匹配近似消息传递算法,并证明梯度下降可达到最优性能。AIM 为理解现代架构中的自注意力层学习提供了可解析的实验场。
原文摘要 · Abstract (English)
We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention models, AIM allows full-width key and query matrices, aligning more closely with practical transformers. Using tools from statistical mechanics and random matrix theory, we derive closed-form predictions for Bayes-optimal generalization error and identify sharp phase transitions as a function of sample complexity, model width, and sequence length. We propose a matching approximate message passing algorithm and show that gradient descent can reach optimal performance. AIM offers a solvable playground for understanding learning in self-attention layers, that are key components of modern architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。