arXiv:2410.23610stat.MLcs.LG2024-10NeurIPS被引 10

证明大模型Transformer训练能全局收敛,为深度学习优化提供理论支撑。

Global Convergence in Training Large-Scale Transformers

  • 构建Transformer的平均场极限,将其优化过程建模为泊松方程。
  • 在小权重衰减下,梯度流可收敛到全局最优解。
  • 提出新分析工具,适用于非均匀、局部光滑的Transformer结构。

尽管Transformer在多个领域广泛应用,但其大规模模型设置下的优化保证仍不明确。本文严格分析了带权重衰减正则化的Transformer梯度流收敛性质。首先,我们构建了大规模Transformer的平均场极限,证明当模型宽度和深度趋于无穷时,梯度流收敛于Wasserstein梯度流,由偏微分方程(PDE)描述。其次,我们证明当权重衰减正则化参数足够小时,梯度流可达到与该PDE解一致的全局最小值。分析基于一系列新颖的平均场技术,这些技术专为Transformer设计。相比现有深度网络工具(Lu等, 2020)要求齐次性和全局Lipschitz光滑性,我们的方法仅假设部分齐次性和局部Lipschitz光滑性,具有更强适用性,相关技术可能具有独立研究价值。

原文摘要 · Abstract (English)

Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in training Transformers with weight decay regularization. First, we construct the mean-field limit of large-scale Transformers, showing that as the model width and depth go to infinity, gradient flow converges to the Wasserstein gradient flow, which is represented by a partial differential equation. Then, we demonstrate that the gradient flow reaches a global minimum consistent with the PDE solution when the weight decay regularization parameter is sufficiently small. Our analysis is based on a series of novel mean-field techniques that adapt to Transformers. Compared with existing tools for deep networks (Lu et al., 2020) that demand homogeneity and global Lipschitz smoothness, we utilize a refined analysis assuming only $\textit{partial homogeneity}$ and $\textit{local Lipschitz smoothness}$. These new techniques may be of independent interest.

Transformer优化理论平均场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。