解析过参数矩阵分解中梯度流的增量学习机制,揭示其低秩近似原理。
Understanding Incremental Learning with Closed-form Solution to Gradient Flow on Overparamerterized Matrix Factorization
- 通过求解黎卡提型微分方程,获得梯度流的闭式解。
- 小初始化下不同奇异值学习时间尺度分离,实现按大小顺序逐个学习。
- 为理解神经网络隐式正则化提供新视角,适合研究优化机制者阅读。
许多关于神经网络的理论研究将其实用性能归因于一阶优化算法在特定初始化下产生的隐式偏差或正则化。以过参数矩阵分解问题为例,在小初始化条件下,梯度流(GF)会随时间依次学习目标矩阵的奇异值,按其大小递减顺序进行。本文针对对称矩阵分解问题,利用求解类似黎卡提矩阵微分方程得到的闭式解,定量分析了梯度流中的增量学习行为。结果表明,该现象源于目标矩阵不同成分学习过程的时间尺度分离;随着初始化规模减小,这种分离愈发显著,从而可实现对目标矩阵的低秩近似。最后,讨论了将此分析拓展至非对称矩阵分解问题的可能性。
原文摘要 · Abstract (English)
Many theoretical studies on neural networks attribute their excellent empirical performance to the implicit bias or regularization induced by first-order optimization algorithms when training networks under certain initialization assumptions. One example is the incremental learning phenomenon in gradient flow (GF) on an overparamerterized matrix factorization problem with small initialization: GF learns a target matrix by sequentially learning its singular values in decreasing order of magnitude over time. In this paper, we develop a quantitative understanding of this incremental learning behavior for GF on the symmetric matrix factorization problem, using its closed-form solution obtained by solving a Riccati-like matrix differential equation. We show that incremental learning emerges from some time-scale separation among dynamics corresponding to learning different components in the target matrix. By decreasing the initialization scale, these time-scale separations become more prominent, allowing one to find low-rank approximations of the target matrix. Lastly, we discuss the possible avenues for extending this analysis to asymmetric matrix factorization problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。