arXiv:2509.11426cs.LGcs.IT2025-09被引 3

揭示非凸梯度下降在高维下的长期轨迹规律,统一解释其泛化性与隐式正则化现象。

Long-time dynamics and universality of nonconvex gradient descent

  • 提出可追踪的确定性向量模型,描述梯度下降在大维度下的演化路径。
  • 证明算法在多种结构链接函数下全局收敛,且对随机初始化具有普适性。
  • 开发无需数据的参数估计算法,可用于超参调优和运行时间预测。

本文建立了一种通用方法,用于刻画广义单指标模型中非凸梯度下降在大纵横比(large aspect ratio)下的长期轨迹行为。在此条件下,我们证明每个迭代步的梯度下降点会集中在称为‘高斯理论梯度下降’的确定性向量附近,其动态可通过两个标量的递推方程系统追踪。该集中性保证对广泛的设计矩阵类成立,并在长时间尺度上持续有效,直至算法收敛或发散。此外,我们的方法揭示了梯度下降迭代点通常近似独立于数据,且与特征向量高度不相干,这一现象此前仅在高斯数据特定模型中被称作‘隐式正则化’。作为理论应用,我们在回归设置中展示了两种不同性质的实例:其一,证明了对一类结构化链接函数,任意独立初始化下的非凸梯度下降全局收敛,并建立了相位恢复中随机初始化梯度下降在大纵横比下的普适性;其二,提出一种无需数据的迭代算法,用于在整个梯度下降轨迹上估计状态演化参数,为超参调优、运行时间判断等实际任务提供低成本且统计有效的工具。作为分析副产品,我们发现大纵横比下,高斯理论梯度下降与近期关于常时域梯度下降的动力学平均场理论结果一致。

原文摘要 · Abstract (English)

This paper develops a general approach to characterize the long-time trajectory behavior of nonconvex gradient descent in generalized single-index models in the large aspect ratio regime. In this regime, we show that for each iteration the gradient descent iterate concentrates around a deterministic vector called the `Gaussian theoretical gradient descent', whose dynamics can be tracked by a state evolution system of two recursive equations for two scalars. Our concentration guarantees hold universally for a broad class of design matrices and remain valid over long time horizons until algorithmic convergence or divergence occurs. Moreover, our approach reveals that gradient descent iterates are in general approximately independent of the data and strongly incoherent with the feature vectors, a phenomenon previously known as the `implicit regularization' effect of gradient descent in specific models under Gaussian data. As an illustration of the utility of our general theory, we present two applications of different natures in the regression setting. In the first, we prove global convergence of nonconvex gradient descent with general independent initialization for a broad class of structured link functions, and establish universality of randomly initialized gradient descent in phase retrieval for large aspect ratios. In the second, we develop a data-free iterative algorithm for estimating state evolution parameters along the entire gradient descent trajectory, thereby providing a low-cost yet statistically valid tool for practical tasks such as hyperparameter tuning and runtime determination. As a by-product of our analysis, we show that in the large aspect ratio regime, the Gaussian theoretical gradient descent coincides with a recent line of dynamical mean-field theory for gradient descent over the constant-time horizon.

非凸优化梯度下降统计学习泛化性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。