用快慢微分方程重看Transformer预训练,揭示深层耦合机制。
Attention is Just Another Name for Coupling? A Fast-Slow ODE Perspective on Hierarchical Pretraining
- 将深度预训练视为带非自治特征的快慢微分流,线性化后为层映射乘积。
- 饱和深度后流动分解为块粗粒化过程,跨块耦合仅存在于可见方向。
- 稳定性约束决定模型能力,数据结构决定实际使用的信息通道。
我们将Transformer预训练重新解释为沿深度方向的快慢奇异摄动流,未绑定权重作为其非自治特征。线性化动力学表现为层映射的深度有序乘积。在单个令牌的参考轨迹上,线性化层可沿冻结注意力核的特征基分解。达到可计算的饱和深度后,流动通过块粗粒化进行分解——即运行各层等价于运行粗变量,双重意义。支持在衰减束上的权重扰动,对标志性轨迹的持久分量和冻结核在首阶无影响,因此参数空间被划分为可见与不可见方向,跨块慢路径耦合完全位于可见侧。慢路径可承载的门限大小受稳定性裕度限制。数据层面:若块输出服从指数族分布,块均值池化即可捕获慢路径可用全部信息;但若相邻块无共享结构,则任何跨块通道无法提升预测,门限幅度在预测风险中不可见。稳定性界定架构可能行为,数据决定其实际表现。
原文摘要 · Abstract (English)
We re-interpret Transformer pretraining as a fast-slow, singularly perturbed flow along depth, with untied weights as its non-autonomous feature. The linearised dynamics is a depth-ordered product of layer maps. Along a token-homogeneous reference trajectory, the linearised layer factorises along the eigenbasis of a frozen attention kernel. Past a computable saturation depth, the flow factors through the block coarse-graining -- in other words, running the layers is running the coarse variable, dually. Weight perturbations supported on the decaying bundle move neither the persistent component of the distinguished trajectory nor the frozen kernel to first order, so the framework partitions parameter space into visible and invisible directions, with the cross-block coupling of the slow path sitting entirely on the visible side. How large a gate the slow path can carry is bounded by a stability margin. On the data side: if block emissions follow an exponential family, block-mean pooling captures all the information the slow path can use; but if neighbouring blocks carry no shared structure, no cross-block channel can help the prediction, and the gate amplitude is invisible in the prediction risk. Stability delimits what the architecture may do; the data decides what it will.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。