揭示大提示下softmax注意力如何趋近线性,为理论分析提供新工具。
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
- 基于测度论框架,证明大提示时softmax退化为线性算子。
- 给出有限提示与无限提示间收敛速度的非渐近界,训练全程稳定。
- 适用于长提示场景的优化分析,可迁移线性注意力结论至softmax。
Softmax注意力是Transformer的核心组件,但其非线性结构给理论分析带来挑战。本文建立统一的测度基框架,研究有限与无限提示下的单层softmax注意力。对于独立同分布高斯输入,我们证明在无限提示极限下,softmax算子收敛为作用于输入标记测度的线性算子。基于此,我们建立了输出与梯度的非渐近集中界,量化了有限提示模型向无限提示版本的逼近速度,并证明该收敛在一般次高斯标记的上下文学习设置中沿整个训练轨迹保持稳定。针对上下文线性回归情形,利用可解析的无限提示动力学,分析了有限提示长度下的训练过程。结果表明,当提示足够长时,softmax注意力继承了线性注意力的分析结构,使专为线性注意力设计的优化分析可直接迁移到softmax注意力,为大提示场景下的软注意力训练动态与统计行为研究提供了原理性且普适的工具包。
原文摘要 · Abstract (English)
Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis. We develop a unified, measure-based framework for studying single-layer softmax attention under both finite and infinite prompts. For i.i.d. Gaussian inputs, we lean on the fact that the softmax operator converges in the infinite-prompt limit to a linear operator acting on the underlying input-token measure. Building on this insight, we establish non-asymptotic concentration bounds for the output and gradient of softmax attention, quantifying how rapidly the finite-prompt model approaches its infinite-prompt counterpart, and prove that this concentration remains stable along the entire training trajectory in general in-context learning settings with sub-Gaussian tokens. In the case of in-context linear regression, we use the tractable infinite-prompt dynamics to analyze training at finite prompt length. Our results allow optimization analyses developed for linear attention to transfer directly to softmax attention when prompts are sufficiently long, showing that large-prompt softmax attention inherits the analytical structure of its linear counterpart. This, in turn, provides a principled and broadly applicable toolkit for studying the training dynamics and statistical behavior of softmax attention layers in large prompt regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。