权重归一化让矩阵感知更快收敛,且越过参数越快。
On the Benefits of Weight Normalization for Overparameterized Matrix Sensing
- 用黎曼优化的权重归一化实现线性收敛
- 收敛速度比传统方法快指数级
- 过参数越多,迭代和样本复杂度越低
尽管归一化技术在深度学习中广泛应用,其理论理解仍相对有限。本文研究了在过参数化矩阵感知问题中应用(广义)权重归一化(WN)的益处。我们证明,结合黎曼优化的权重归一化可实现线性收敛,相比不使用权重归一化的标准方法,收敛速度呈指数级提升。分析进一步表明,随着过参数化程度的增加,迭代复杂度和样本复杂度均呈多项式改善。据我们所知,这是首个对权重归一化如何利用过参数化加速矩阵感知收敛的刻画。
原文摘要 · Abstract (English)
While normalization techniques are widely used in deep learning, their theoretical understanding remains relatively limited. In this work, we establish the benefits of (generalized) weight normalization (WN) applied to the overparameterized matrix sensing problem. We prove that WN with Riemannian optimization achieves linear convergence, yielding an exponential speedup over standard methods that do not use WN. Our analysis further demonstrates that both iteration and sample complexity improve polynomially as the level of overparameterization increases. To the best of our knowledge, this work provides the first characterization of how WN leverages overparameterization for faster convergence in matrix sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。