arXiv:2509.21936cs.LGcond-mat.dis-nn2025-09中稿 · ICLR被引 7

揭示了注意力机制中Softmax为何优于线性注意力的统计优势

Statistical Advantage of Softmax Attention: Insights from Single-Location Regression

  • 用统计物理方法分析高维下注意力预测器的泛化性能
  • 证明Softmax在总体情况下达到贝叶斯最优,线性注意力则无法做到
  • 适用于理解大模型注意力机制原理的研究者与理论分析人员

大型语言模型依赖带有Softmax激活的注意力机制。然而,Softmax为何优于其他替代方案(如逐分量或线性)仍缺乏清晰理解,许多理论研究聚焦于更易分析的线性化注意力。本文通过单位置回归任务,系统研究注意力机制的统计特性:输出仅依赖于随机位置输入标记的线性变换。基于统计物理思想,我们建立了高维极限下的注意力预测器分析框架,其中泛化性能由少量序参数刻画。在总体层面,我们证明Softmax可达到贝叶斯风险,而线性注意力存在根本性缺陷。进一步分析其他激活函数,识别出最优性能所需的关键性质。最后,在有限样本情形下,我们给出了测试误差的渐近刻画:尽管Softmax不再为贝叶斯最优,但仍持续优于线性注意力。讨论了其与梯度优化算法的关联。

原文摘要 · Abstract (English)

Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly understood, and many theoretical works have focused on the easier-to-analyze linearized attention. In this work, we address this gap through a principled study of the single-location regression task, where the output depends on a linear transformation of a single input token at a random location. Building on ideas from statistical physics, we develop an analysis of attention-based predictors in the high-dimensional limit, where generalization performance is captured by a small set of order parameters. At the population level, we show that softmax achieves the Bayes risk, whereas linear attention fundamentally falls short. We then examine other activation functions to identify which properties are necessary for optimal performance. Finally, we analyze the finite-sample regime: we provide an asymptotic characterization of the test error and show that, while softmax is no longer Bayes-optimal, it consistently outperforms linear attention. We discuss the connection with optimization by gradient-based algorithms.

注意力机制理论分析统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。