arXiv:2506.18656stat.MLcs.LG2025-06

对比非线性注意力与线性回归的插值误差,发现注意力在随机输入下误差更大

On the Interpolation Error of Nonlinear Attention versus Linear Regression

  • 在高维下分析注意力插值误差,推导出均方误差的极限表达式
  • 随机输入时注意力误差高于线性回归,但有结构信号时可反超
  • 理论结合实验,适合关注注意力机制原理的研究者

注意力已成为现代机器学习中捕捉输入标记长程依赖的核心组件,其固有的并行结构支持数据和模型参数规模快速增加下的高效扩展。尽管作用关键,对注意力的理论理解,尤其是非线性情形,进展相对缓慢。本文在输入标记数n与嵌入维度p均大且相近的高维条件下,对非线性注意力的插值误差进行了精确刻画。在信号加噪声的数据模型下,对于固定注意力权重,我们推导出均方插值误差的显式(极限)表达式。借助随机矩阵论的最新进展,表明非线性注意力在随机输入上通常比线性回归具有更大的插值误差。然而,当输入包含结构化信号,特别是注意力权重与信号方向对齐时,这一差距会消失甚至逆转。理论结果得到数值实验的支持。

原文摘要 · Abstract (English)

Attention has become the core building block of modern machine learning (ML) by efficiently capturing the long-range dependencies among input tokens. Its inherently parallelizable structure allows for efficient performance scaling with the rapidly increasing size of both data and model parameters. Despite its central role, the theoretical understanding of Attention, especially in the nonlinear setting, is progressing at a more modest pace. This paper provides a precise characterization of the interpolation error for a nonlinear Attention, in the high-dimensional regime where the number of input tokens $n$ and the embedding dimension $p$ are both large and comparable. Under a signal-plus-noise data model and for fixed Attention weights, we derive explicit (limiting) expressions for the mean-squared interpolation error. Leveraging recent advances in random matrix theory, we show that nonlinear Attention generally incurs a larger interpolation error than linear regression on random inputs. However, this gap vanishes, and can even be reversed, when the input contains a structured signal, particularly if the Attention weights align with the signal direction. Our theoretical insights are supported by numerical experiments.

注意力机制理论分析高维统计插值误差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。