大步长梯度下降让深度线性网络重获路径对称性
Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway

- 用大步长离散梯度下降,发现路径间信号会重新分配
- 深层网络中多路径分布比单路径更平坦,降低尖锐极小值
- 适合研究深度网络训练动态与对称性恢复的学者
近期对多路径深度线性网络的分析采用梯度流(Gradient Flow)预测‘赢家通吃’现象:路径对称性被打破,每个特征集中于单一路径。本文表明,使用大步长的离散梯度下降(GD)则呈现不同结果。我们证明,单路径解为尖锐极小值,而将信号分布在多路径上可使尖锐度降低,降幅随路径数和网络深度增加而增大。因此,尽管早期训练仍体现深度引发的对称性破缺,但处于稳定性边缘的振荡最终覆盖此趋势,促使网络进入重平衡阶段,信号在路径间重新分配。上述结果揭示了深度如何影响路径竞争机制,并解释为何大步长GD倾向于共享表示而非持久的单路径主导。
原文摘要 · Abstract (English)
Recent analyses of multi-pathway Deep Linear Networks use Gradient Flow to predict a "winner-takes-all" specialization in which path symmetry breaks and each feature concentrates in a single pathway. In this work, we show that discrete Gradient Descent (GD) with a large step size tells a different story. We prove that single-path solutions are sharp minima, whereas distributing signals across pathways reduces sharpness by a factor that decreases with both the number of pathways and depth. Consequently, while early training reproduces the depth-driven symmetry breaking predicted by GF, oscillations at the Edge of Stability subsequently override this tendency and drive the network into a re-balancing phase, where signals redistribute across pathways. Together, these results clarify how depth shapes pathway competition and explain why large-step GD favors shared representations rather than persistent single-pathway dominance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。