揭示长序列注意力的临界规律,解释为何模型会聚焦或平均。
Scaling Limits of Long-Context Transformers

- 通过分析查询与随机键的几何关系,揭示注意力选择性的临界条件。
- 当温度参数满足 β_n ∗ ∼ n^(2/(d−1)) 时,注意力开始聚焦最近键。
- 发现子临界状态下注意力近似实现反向热方程,适合理解模型演化机制。
我们研究了固定查询下,基于球面上独立同分布键的softmax自注意力在长上下文极限中的行为,将逆温度β_n作为调控参数,决定注意力是否退化为均匀平均或坍缩至最近单键。研究发现,选择性出现的临界尺度由查询附近距离分布的局部指数决定,而非全局上下文特征,对于球面S^{d−1}上的均匀键,临界尺度为β_n^∗ ∼ n^(2/(d−1))。此外,我们刻画了所有β_n regimes下有序注意力权重和注意力输出的极限分布:亚临界态下输出退化为查询附近的局部平均,具有确定性偏差与高斯波动;临界态下若干最近键共同承载宏观质量,未发生单键坍缩;超临界态下所有质量集中于最近键。值得注意的是,在恒等值矩阵的亚临界情况下,注意力映射近似实现反向热方程。
原文摘要 · Abstract (English)
We study the long-context limit of softmax self-attention with a fixed query and a random context of $n$ i.i.d. keys on the sphere, viewing the inverse temperature $β_n$ as the scaling parameter that decides whether attention degenerates into uniform averaging or collapses onto the single closest key. We show that the critical scale at which selectivity emerges is determined by the local exponent of the distance-to-query distribution near zero rather than by global features of the context, and scales like $β_n^\ast \asymp n^{2/(d-1)}$ for uniform keys on $\mathbb{S}^{d-1}$. Furthermore, we characterize the limiting laws of the ordered attention weights and of the attention output across all regimes of $β_n$: a subcritical regime in which the output reduces to a local average around $q$ with explicit deterministic bias and Gaussian fluctuations; a critical regime in which a finite collection of nearest keys retains macroscopic mass without single-key collapse; and a supercritical regime in which all mass concentrates on the closest key. Of notable interest is the subcritical case with identity value matrix where the attention map approximately implements a backward heat equation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。