训练让注意力聚集的分组状态逃逸,揭示了模型动态新机制。
Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers
- 构建带噪声的均场Transformer模型,仅线性前馈层受正则化训练
- 训练后期token分布从聚集态脱离,尤其在深层网络中出现相变
- 提出熵正则交互能模型,为训练-推理一体化建模提供新框架
Transformer通过逐层迭代变换标记表示进行推理。这种层间计算已得到实证研究,近期的均场理论解释了注意力如何使标记分布趋向聚类。然而,现有均场分析多将参数视为给定,未阐明训练如何重塑这一聚类图景。本文研究一个带噪声的均场Transformer,其中仅有参数线性的前馈网络(FFN)在$ L^2 $正则化下被训练。我们发现并分析了一种由训练引发的动力学相变:初始阶段跟随注意力驱动的聚类,但在接近最终层时,标记分布可脱离聚类区域。数学分析基于熵正则化的相互作用能量,捕捉了注意力的聚类偏置。更广泛地,我们的结果指向一种训练感知的均场理论,将训练与推理动态统一建模。
原文摘要 · Abstract (English)
Transformers perform inference by iteratively transforming token representations across layers. This layerwise computation has been studied empirically, and recent mean-field theories of Transformer dynamics explain how attention can drive token distributions toward clustering. However, existing mean-field analyses largely treat model parameters as prescribed, leaving open how training reshapes this clustering picture. We study this question in a noisy mean-field Transformer in which only a parameter-linear FFN is trained under $L^2$ regularization. We find and analyze a training-induced phase in the dynamics: after initially following attention-driven clustering, the token distribution can leave the clustered regime near the final layers. Our mathematical analysis is based on an entropy-regularized interaction energy that captures the clustering bias of attention. More broadly, our results point toward a training-aware mean-field theory of Transformer dynamics, in which training and inference dynamics are treated together.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。