提出无需过参数化的损失曲面刻画方法,保障梯度优化收敛性。
Loss Landscape Characterization of Neural Networks without Over-Parametrization
- 引入新函数类刻画深度模型损失曲面,无需过度过参数化。
- 证明在该假设下梯度优化器具有收敛理论保证。
- 适用于多种深度学习模型,兼具理论严谨与实证支持。
优化方法在现代机器学习中起关键作用,驱动了深度学习模型的显著实证成就。这些成功尤为突出,因为模型损失曲面具有复杂的非凸特性。然而,确保优化方法收敛需特定目标函数结构条件,实践中极少满足。一个典型例子是近年来备受关注的Polyak-Lojasiewicz(PL)不等式。但验证深度神经网络满足此类假设通常需要大量且往往不切实际的过参数化。为克服这一局限,我们提出一类新函数,可刻画现代深度模型的损失曲面而无需大规模过参数化,并能包含鞍点。关键的是,我们证明在该假设下,基于梯度的优化器具备理论收敛保证。最后,通过理论分析和跨多种深度学习模型的实证实验,验证了该函数类的有效性。
原文摘要 · Abstract (English)
Optimization methods play a crucial role in modern machine learning, powering the remarkable empirical achievements of deep learning models. These successes are even more remarkable given the complex non-convex nature of the loss landscape of these models. Yet, ensuring the convergence of optimization methods requires specific structural conditions on the objective function that are rarely satisfied in practice. One prominent example is the widely recognized Polyak-Lojasiewicz (PL) inequality, which has gained considerable attention in recent years. However, validating such assumptions for deep neural networks entails substantial and often impractical levels of over-parametrization. In order to address this limitation, we propose a novel class of functions that can characterize the loss landscape of modern deep models without requiring extensive over-parametrization and can also include saddle points. Crucially, we prove that gradient-based optimizers possess theoretical guarantees of convergence under this assumption. Finally, we validate the soundness of our new function class through both theoretical analysis and empirical experimentation across a diverse range of deep learning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。