证明了四层矩阵分解在随机初始化下梯度下降的全局收敛性。
Global Convergence of Four-Layer Matrix Factorization under Random Initialization
- 引入新方法分析梯度下降避免鞍点的特性
- 在目标矩阵满足条件时实现多项式时间收敛
- 适合研究深度学习理论与优化机制的读者
深度矩阵分解的梯度下降动态被广泛视为深度神经网络的简化理论模型。尽管两层矩阵分解的收敛理论已较为完善,但至今尚未建立一般深度矩阵分解在随机初始化下的全局收敛保证。为填补这一空白,本文在目标矩阵满足特定条件且采用标准平衡正则化项的前提下,给出了四层矩阵分解中随机初始化梯度下降的多项式时间全局收敛保证。分析中引入新技巧,揭示了梯度下降动态的鞍点规避性质,并将前期理论扩展至刻画各层权重特征值的变化过程。
原文摘要 · Abstract (English)
Gradient descent dynamics on the deep matrix factorization problem is extensively studied as a simplified theoretical model for deep neural networks. Although the convergence theory for two-layer matrix factorization is well-established, no global convergence guarantee for general deep matrix factorization under random initialization has been established to date. To address this gap, we provide a polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization, given certain conditions on the target matrix and a standard balanced regularization term. Our analysis employs new techniques to show saddle-avoidance properties of gradient decent dynamics, and extends previous theories to characterize the change in eigenvalues of layer weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。