揭示Transformer在高斯输入下的动态可到达性与渐近行为
Reachability and asymptotics of Gaussian Transformer dynamics
- 将Transformer视为概率测度上的非线性控制系统,解析其高斯分布保持性
- 时间变化控制下可精确达到同秩目标高斯分布,时间不变时出现稳定或爆炸性演化
- 理论与实验结合,适用于理解大模型中间层输出的统计特性
我们将大规模语言模型核心架构Transformer的数据传播建模为概率测度空间上的非线性控制系统。针对具有自注意力和仿射前馈层的均场Transformer模型,证明了高斯分布沿其诱导流始终保持高斯性。该不变性将无限维测度动力学降维为有限维双线性控制系统,描述均值与协方差的演化,将Transformer的表达能力转化为指定高斯矩的可达性问题,并揭示其与经典滤波与控制中的Riccati型方程的新关联。对于时变控制,我们证明了任何目标高斯分布若协方差矩阵与初始同秩,则可在有限时间内精确到达——此秩约束是动力学的内在不变量。对于时不变参数,推导出显式谱条件:要么渐近趋于正定平衡点,要么协方差在有限时间内爆炸。数值实验验证了理论:实际Transformer在高斯输入下,早期与中期层仍保持与矩匹配的高斯分布;固定注意力矩阵的模型再现预测的协方差演化模式:稳定配置下有界演化,不稳定配置下发生爆炸。
原文摘要 · Abstract (English)
We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system on the space of probability measures. For the mean-field Transformer model with self-attention and affine feed-forward layers, we prove that Gaussian distributions remain exactly Gaussian along the induced flow. This invariance reduces the infinite-dimensional measure dynamics to a finite-dimensional bilinear control system governing the evolution of the mean and covariance, reformulates the expressive capacity of Transformers as a reachability problem for prescribed Gaussian moments, and reveals a novel connection with Riccati-type equations from classical filtering and control. For time-varying controls, we prove exact finite-time reachability of any target Gaussian distribution whose covariance matrix has the same rank as the initial one, this rank constraint being an intrinsic invariant of the dynamics. For time-invariant parameters, we derive explicit spectral conditions leading either to asymptotic stability toward positive-definite equilibria or to finite-time blow-up of the covariance. Numerical experiments complement the theory by showing that practical Transformers with Gaussian inputs remain close to moment-matched Gaussian distributions through early and intermediate layers, while Transformers with prescribed attention matrices reproduce the predicted covariance regimes: bounded evolution in stabilizing configurations and blow-up in destabilizing ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。