arXiv:2509.25351cs.LGstat.ML2025-09被引 5

大步长下梯度下降陷入混沌,收敛区域呈分形结构。

Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region

  • 大步长导致参数空间出现分形结构,收敛依赖初始值敏感性。
  • 临界步长附近,初始化微小变化可引发完全不同的收敛结果。
  • 正则化加剧敏感性,使收敛与发散边界呈现复杂分形特征。

我们研究矩阵分解中的梯度下降,发现大步长下参数空间会形成分形结构。在标量-向量分解中,推导出精确的临界步长,并表明在临界附近,最小化器的选择对初始化极为敏感。此外,添加正则化会放大这种敏感性,导致收敛与发散的初始条件之间形成分形边界。该分析扩展至正交初始化下的通用矩阵分解。结果揭示,在近临界步长时,梯度下降进入混沌状态,训练结果不可预测,且不存在简单的隐式偏好,如平衡性、最小范数或平坦性。

原文摘要 · Abstract (English)

We examine gradient descent in matrix factorization and show that under large step sizes the parameter space develops a fractal structure. We derive the exact critical step size for convergence in scalar-vector factorization and show that near criticality the selected minimizer depends sensitively on the initialization. Moreover, we show that adding regularization amplifies this sensitivity, generating a fractal boundary between initializations that converge and those that diverge. The analysis extends to general matrix factorization with orthogonal initialization. Our findings reveal that near-critical step sizes induce a chaotic regime of gradient descent where the training outcome is unpredictable and there are no simple implicit biases, such as towards balancedness, minimum norm, or flatness.

优化理论分形结构梯度下降混沌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。