arXiv:2606.22406cs.LGcs.IT2026-06

证明了注意力模型能从噪声中精准提取信号方向。

Asymptotic Signal Subspace Recovery in Softmax Attention Models

论文配图:Asymptotic Signal Subspace Recovery in Softmax Attention Models
图 1 · 摘自论文原文
  • 基于随机梯度上升的注意力学习,最终聚焦于信号子空间。
  • 高维下查询向量几乎必然收敛到信号方向。
  • 为注意力机制在噪声环境中的信息提取提供理论支撑。

注意力机制在从大量标记中识别相关信息方面表现出显著的实证成功,但其背后的理论原理仍不清晰。本文研究一种简化的softmax注意力模型,其中查询向量通过随机梯度上升从信息性与干扰性标记中学习。利用模型对称性,推导出总体目标函数,并刻画了决定学习动态的极限常微分方程。结合随机逼近与动力系统理论,建立了随机学习算法与其确定性极限之间的严格联系。主要结果表明,在合适的高维尺度假设和标准步长条件下,学习到的查询向量几乎必然收敛至由潜在信息方向张成的一维信号子空间。等价地,查询渐近地恢复了潜在信号,仅存在固有的符号歧义。这些结果为理解注意力机制作为高维噪声环境中信号提取手段提供了严格的理论基础,并从动力系统视角揭示了注意力如何在强噪声下发现相关信息。

原文摘要 · Abstract (English)

Attention mechanisms have demonstrated remarkable empirical success in identifying relevant information from large collections of tokens, yet the theoretical principles underlying this behavior remain poorly understood. We study a stylized softmax-attention model in which a query vector is learned by stochastic gradient ascent from a collection of informative and nuisance tokens. Exploiting the symmetry of the model, we derive a population objective and characterize the limiting ordinary differential equation governing the learning dynamics. Using tools from stochastic approximation and dynamical systems theory, we establish a rigorous connection between the stochastic learning algorithm and its deterministic limit. Our main result shows that, under suitable high-dimensional scaling assumptions and standard step-size conditions, the learned query converges almost surely to the one-dimensional signal subspace spanned by the latent informative direction. Equivalently, the query asymptotically recovers the latent signal up to the intrinsic sign ambiguity. These results provide a rigorous theoretical foundation for understanding attention mechanisms as signal extraction procedures in high-dimensional noisy environments and offer a dynamical-systems perspective on how attention discovers relevant information in the presence of substantial noise.

注意力机制信号提取理论分析高维学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。