用最优控制理论分析Transformer泛化能力,建立可量化边界。
Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness
- 将Transformer训练视为马尔可夫控制问题,基于测度空间建模动态。
- 在有限样本下推导出显式泛化界,误差由量化近似决定。
- 连接泛化与Wasserstein分布鲁棒优化,适合理论研究者参考。
我们为采用动态规划递推训练的Transformer推导了有限样本泛化界。基于双重提升的测度值变压器动力学表述,我们将数据集视为输入输出测度对上的概率律,从而将训练问题视作有限时域马尔可夫控制问题。随后分析一个通过量化状态、动作及测度状态空间得到的量化模型,并利用有限度量空间上经验律的浓度不等式以及价值函数的Lipschitz稳定性估计,推导出明确的有限样本泛化界。这些界通过显式的近似误差转移至原始模型。最后,我们证明相同方法可导出训练问题的分布鲁棒控制形式,将Transformer泛化性与Wasserstein分布鲁棒优化相联系。
原文摘要 · Abstract (English)
We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem. We then analyze a quantized model, derived by quantizing the state, action, and measure-state spaces, and derive explicit finite-sample generalization bounds using concentration inequalities for empirical laws on finite metric spaces together with a Lipschitz stability estimate for the value function. These bounds are transferred to the base model at the cost of an explicit approximation error. Finally, we show that the same machinery yields a distributionally robust control formulation of the training problem, connecting Transformer generalization to Wasserstein distributionally robust optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。