揭示注意力模型学习成对交互的最优收敛速率,与维度无关。
Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
- 通过权重矩阵与非线性激活函数建模令牌间交互
- 收敛速率达 $M^{-\frac{2β}{2β+1}}$,与维度无关
- 适用于理解注意力机制的理论基础与训练指导
我们研究单层注意力类模型中学习成对交互的收敛速率,其中令牌通过权重矩阵和非线性激活函数相互作用。证明了最小最大速率为 $M^{-\frac{2β}{2β+1}}$,其中 $M$ 为样本量,$β$ 为激活函数的 Hölder 光滑度。重要的是,该速率在满足 $rd \le (M/\log M)^{\frac{1}{2β+1}}$ 的条件下,与嵌入维度 $d$、令牌数 $N$ 及权重矩阵秩 $r$ 无关。结果揭示了注意力类模型在统计上的根本高效性,即使权重矩阵与激活函数不可单独识别,也为理解注意力机制提供了理论依据,并指导实际训练。
原文摘要 · Abstract (English)
We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation function. We prove that the minimax rate is $M^{-\frac{2β}{2β+1}}$, where $M$ is the sample size and $β$ is the Hölder smoothness of the activation function. Importantly, this rate is independent of the embedding dimension $d$, the number of tokens $N$, and the rank $r$ of the weight matrix, provided that $rd \le (M/\log M)^{\frac{1}{2β+1}}$. These results highlight a fundamental statistical efficiency of attention-style models, even when the weight matrix and activation are not separately identifiable, and provide a theoretical understanding of attention mechanisms and guidance on training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。