arXiv:2605.16747cs.LGmath.AP2026-05被引 1

揭示大上下文下Transformer的统计演化规律,给出误差收敛理论。

Propagation of Chaos in Contextual Flow Maps

  • 用上下文流映射建模Transformer,将上下文长度视为统计资源。
  • 证明有限与无限上下文模型间误差以n^{-1/d}或n^{-1/2}速率收敛。
  • 适用于Transformer等模型,为训练稳定性提供新分析工具。

我们通过上下文流映射(CFMs)这一抽象框架,发展了大上下文场景下Transformer的定量统计理论:在注意力层堆叠中,一个特殊标记随上下文测度动态演化。有限上下文模型近似于理想化的无限上下文系统,其中上下文测度被其总体分布取代,使上下文长度 $n$ 成为统计资源。利用动力学中的McKean--Vlasov结构和传播混沌的经典方法,我们建立了前向界,统一控制沿深度的有限-无限上下文CFM偏差;以及后向界,统一控制在线梯度下降迭代中对应训练轨迹的偏差。两个界对一般CFMs达到最优的Wasserstein率 $n^{-1/d}$,对包含Transformer在内的受限类模型达到参数率 $n^{-1/2}$。分析依赖于损失梯度的新欧拉伴随形式及由此产生的前向-伴随系统的稳定性估计,两者可能具有独立研究价值。

原文摘要 · Abstract (English)

We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a distinguished token in the presence of a contextual measure across a stack of attention blocks. Within this framework, the finite-context model approximates an idealized infinite-context system in which the contextual measure is replaced by its underlying population, so that the context length $n$ becomes a statistical resource. Exploiting the McKean--Vlasov structure of the dynamics and the classical machinery of propagation of chaos, we establish a forward bound controlling the deviation between the finite- and infinite-context CFMs uniformly along depth, and a backward bound controlling the deviation between the corresponding training trajectories uniformly across iterations of online gradient descent. Both bounds achieve the optimal Wasserstein rate $n^{-1/d}$ for general CFMs and parametric rate $n^{-1/2}$ for a restricted class of CFMs that includes transformers as a special case. The analysis rests on a new Eulerian adjoint formulation of the loss gradient and stability estimates for the resulting forward--adjoint system, both of which may be of independent interest.

Transformer统计力学扩散模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。