揭示了自注意力中温度缩放的统一理论,解决长期上下文稳定性难题。
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
- 基于每行注意力的间隙计数函数 $N_n$ 推导出临界温度尺度
- 低于该尺度时顶级注意力项无法分离,高于则熵崩溃
- 可诊断从理论模型到实际Transformer的注意力得分行为
长度相关对数归一化广泛用于稳定长序列自注意力,但现有分析和方法对上下文长度 $n$ 的逆温度标度存在冲突,范围从 $(\log n)^{1/2}$ 到 $\log n$ 及 $(\log n)^2$。本文提出通用理论,表明理想标度取决于每行注意力的间隙计数函数 $N_n$:统计每个间隙中位于最大值之下的竞争者数量,定义上尾累积尺度,并证明该尺度即为softmax集中性的临界逆温度尺度——低于此尺度时,顶级竞争者保持未分离;高于时,注意力熵发生坍塌。该框架统一了此前各类标度律,适用于从理想理论模型到实际Transformer的注意力得分族。
原文摘要 · Abstract (English)
Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting inverse-temperature laws for the context length $n$, ranging from $(\log n)^{1/2}$ to $\log n$ and $(\log n)^2$. We provide a general theory showing that the desirable scale is determined by the gap-counting function $N_n$ of each attention row. Counting how many competitors lie within each gap from the maximum, we define an upper-tail accumulation scale and prove that it gives the critical inverse-temperature scale for softmax concentration: below this scale, the top competitors remain unseparated, whereas above it, the attention entropy collapses. This framework unifies prior scaling laws as different $N_n$ and yields a direct diagnostic for attention-score families, from idealized theoretical models to more practical transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。