arXiv:2603.09571cs.LGmath.OC2026-03被引 3

用最优控制理论为Transformer训练提供全局最优解,不依赖梯度。

An Optimal Control Approach To Transformer Training

  • 将Transformer建模为带共享动作的粒子系统,通过概率测度升维实现马尔可夫决策过程。
  • 提出三重量化训练法,证明其近似原问题的最优解,且策略与输入无关。
  • 适用于追求理论严谨性、不依赖梯度优化的场景,如高可靠性模型训练。

本文提出一种严格的最优控制理论方法来训练Transformer,兼顾执行时的输入无关性、问题的集成控制特性及位置依赖性。将Transformer架构建模为具有共享动作的离散时间受控粒子系统,呈现无噪声的McKean-Vlasov动力学。尽管该动力学非马尔可夫,但通过将其提升至概率测度空间,可获得完全可观测的马尔可夫决策过程(MDP)。位置编码被纳入状态空间以保持序列顺序。利用动态规划原理,在温和紧致性假设下证明全局最优策略的存在性。进一步证明,升维后的闭环策略等价于初始分布相关的开环策略,具备输入无关性并兼容标准Transformer训练。为此,我们提出对升维后的MDP进行三重量化——状态空间、概率测度空间和动作空间,并证明任意最优策略对原始训练问题均为近似最优。最后,通过证明价值函数对初始经验测度扰动的连续性以及数据量增大时策略的收敛性,建立了该模型的稳定性和经验一致性。该方法无需光滑性或凸性假设,为基于梯度的训练提供了全局最优且鲁棒的替代方案。

原文摘要 · Abstract (English)

In this paper, we develop a rigorous optimal control-theoretic approach to Transformer training that respects key structural constraints such as (i) realized-input-independence during execution, (ii) the ensemble control nature of the problem, and (iii) positional dependence. We model the Transformer architecture as a discrete-time controlled particle system with shared actions, exhibiting noise-free McKean-Vlasov dynamics. While the resulting dynamics is not Markovian, we show that lifting it to probability measures produces a fully-observed Markov decision process (MDP). Positional encodings are incorporated into the state space to preserve the sequence order under lifting. Using the dynamic programming principle, we establish the existence of globally optimal policies under mild assumptions of compactness. We further prove that closed-loop policies in the lifted is equivalent to an initial-distribution dependent open-loop policy, which are realized-input-independent and compatible with standard Transformer training. To train a Transformer, we propose a triply quantized training procedure for the lifted MDP by quantizing the state space, the space of probability measures, and the action space, and show that any optimal policy for the triply quantized model is near-optimal for the original training problem. Finally, we establish stability and empirical consistency properties of the lifted model by showing that the value function is continuous with respect to the perturbations of the initial empirical measures and convergence of policies as the data size increases. This approach provides a globally optimal and robust alternative to gradient-based training without requiring smoothness or convexity.

Transformer最优控制全局最优理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。