arXiv:2511.02258stat.MLcs.LG2025-11被引 1

揭示高维单层网络中SGD的随机波动规律,指出确定性极限的局限性。

Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks

  • 在临界步长尺度下,引入修正项改变学习动力学相图。
  • 在固定点附近,相关性的扩散极限为均值回复的奥恩斯坦-乌伦贝克过程。
  • 信息指数≥3时系统稳定,指数=2时稳定性取决于步长与噪声水平。

本文研究了在线随机梯度下降(SGD)在高维情况下的渐近行为。基于Ben Arous、Gheissari和Jagannath关于SGD有效动力学的工作,我们分析了单层网络中步长的临界缩放尺度。低于该尺度时,有效动力学由确定性(类弹道)极限主导;在临界尺度下,修正项出现并改变相图。在这些动力学的不动点附近,归一化相关性的扩散极限是一个奥恩斯坦-乌伦贝克过程。具体而言,当信息指数至少为3时,该过程具有均值回复特性;当信息指数为2时,漂移项无普适符号,不动点可能变为排斥型。我们在相位恢复问题中明确展示了这一点,其符号由步长和噪声水平决定。这些结果说明确定性缩放极限无法充分捕捉高维学习动态中的随机波动。

原文摘要 · Abstract (English)

This paper studies the high-dimensional scaling limits of online stochastic gradient descent (SGD). Building on the work of Ben Arous, Gheissari, and Jagannath on the effective dynamics of SGD, we study the critical scaling regime of the step size for single-layer networks. Below this regime the effective dynamics are governed by deterministic (ballistic) limits, whereas at the critical scale a correction term emerges that changes the phase diagram. Near the fixed points of these dynamics, we show that the diffusive (SDE) limit of the rescaled correlation is an Ornstein-Uhlenbeck process. More precisely, it is mean-reverting whenever the information exponent is at least three. At information exponent two the drift has no universal sign, and the fixed point may become repelling; we show this explicitly for phase retrieval, where the sign is determined by the step size and the noise level. These results illustrate the limitations of deterministic scaling limits in capturing stochastic fluctuations in high-dimensional learning dynamics.

随机梯度高维学习动力学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。