深度越深,矩阵补全越倾向低秩解,且避免训练后性能下降。
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
- 发现深度网络的耦合动态是低秩偏置的关键机制。
- 深度≥3时,除非对角初始化,否则必然产生耦合动态并收敛到秩1。
- 解释了为何深度模型能避免预训练后的性能退化现象。
我们通过深度矩阵分解(即深度线性神经网络)研究矩阵补全问题,以考察网络深度对训练动态的影响。尽管该问题简单且重要,但现有理论主要集中在浅层(深度2)模型,无法完全解释深层网络中观察到的隐式低秩偏置。我们识别出耦合动态是这一偏置的核心机制,并证明其随深度增加而增强。在块对角观测下基于梯度流分析,证明:(a) 深度≥3的网络除非对角初始化,否则必出现耦合;(b) 只有在耦合动态下才会收敛至秩1——解决了Menon(2024)针对一类初始化的开放问题。我们还重新审视了矩阵补全中的塑性丧失现象(Kleinman等,2024),即先用少量观测预训练再恢复更多数据反而表现更差。结果表明,深度模型因具备低秩偏置而避免了塑性丧失;而深度2网络在解耦动态下预训练后即使后续训练满足耦合条件也无法收敛至低秩——揭示了该现象的内在机制。
原文摘要 · Abstract (English)
We study matrix completion via deep matrix factorization (a.k.a. deep linear neural networks) as a simplified testbed to examine how network depth influences training dynamics. Despite the simplicity and importance of the problem, prior theory largely focuses on shallow (depth-2) models and does not fully explain the implicit low-rank bias observed in deeper networks. We identify coupled dynamics as a key mechanism behind this bias and show that it intensifies with increasing depth. Focusing on gradient flow under block-diagonal observations, we prove: (a) networks of depth $\geq 3$ exhibit coupling unless initialized diagonally, and (b) convergence to rank-1 occurs if and only if the dynamics is coupled -- resolving an open question by Menon (2024) for a family of initializations. We also revisit the loss of plasticity phenomenon in matrix completion (Kleinman et al., 2024), where pre-training on few observations and resuming with more degrades performance. We show that deep models avoid plasticity loss due to their low-rank bias, whereas depth-2 networks pre-trained under decoupled dynamics fail to converge to low-rank, even when resumed training (with additional data) satisfies the coupling condition -- shedding light on the mechanism behind this phenomenon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。