解析深度对角线网络的优化机制,揭示梯度方法为何在复杂模型中仍有效。
Optimization Insights into Deep Diagonal Linear Networks
- 用重参数化构造可保持凸性的有效参数空间,简化复杂优化问题。
- 证明梯度流在层参数上诱导镜面流动态,实现指数级损失下降。
- 揭示参数化与初始化如何影响训练速度,适合研究优化理论者阅读。
基于梯度的方法在实践中能成功训练高度过参数化的模型,尽管其优化问题显著非凸。理解这些方法为何有效的机制已成为现代优化的核心问题。为在可处理的设置下研究此问题,我们考察深度对角线线性网络。这类多层架构通过重参数化使有效参数保持凸性,同时在优化景观中引入非平凡几何结构。在温和初始化条件下,我们证明层参数上的梯度流在有效参数空间中诱导镜面流动态。这一结构洞察带来了明确的收敛保证,包括在Polyak-Lojasiewicz条件下损失的指数衰减,并阐明了参数化与初始化尺度如何调控训练速度。总体而言,我们的结果表明,尽管深度对角线过参数化看似复杂,但仍可赋予标准梯度方法良好且可解释的优化动态。
原文摘要 · Abstract (English)
Gradient-based methods successfully train highly overparameterized models in practice, even though the associated optimization problems are markedly nonconvex. Understanding the mechanisms that make such methods effective has become a central problem in modern optimization. To investigate this question in a tractable setting, we study Deep Diagonal Linear Networks. These are multilayer architectures with a reparameterization that preserves convexity in the effective parameter, while inducing a nontrivial geometry in the optimization landscape. Under mild initialization conditions, we show that gradient flow on the layer parameters induces a mirror-flow dynamic in the effective parameter space. This structural insight yields explicit convergence guarantees, including exponential decay of the loss under a Polyak-Lojasiewicz condition, and clarifies how the parametrization and initialization scale govern the training speed. Overall, our results demonstrate that deep diagonal over parameterizations, despite their apparent complexity, can endow standard gradient methods with well-behaved and interpretable optimization dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。