揭示随机梯度下降轨迹的长期波动规律,适用于非光滑问题。
Functional Central Limit Theorem for Stochastic Gradient Descent
- 通过尺度变换分析梯度下降路径的渐近形状
- 证明轨迹在长时间下服从扩散极限,描述波动时序结构
- 适合研究优化算法稳定性或非光滑优化的读者
我们研究了在凸目标函数上应用随机梯度下降算法时,其轨迹的渐近形状。在较弱的正则性假设下,我们证明了经过适当尺度变换后的轨迹满足泛函中心极限定理。该结果通过提供轨迹的扩散极限,刻画了算法在最小值附近的长期波动特性。与经典中心极限定理(针对最后迭代点或Polyak-Ruppert平均)不同,这一泛函结果捕捉了波动的时间结构,并适用于非光滑情形,如鲁棒位置估计,包括几何中位数。
原文摘要 · Abstract (English)
We study the asymptotic shape of the trajectory of the stochastic gradient descent algorithm applied to a convex objective function. Under mild regularity assumptions, we prove a functional central limit theorem for the properly rescaled trajectory. Our result characterizes the long-term fluctuations of the algorithm around the minimizer by providing a diffusion limit for the trajectory. In contrast with classical central limit theorems for the last iterate or Polyak-Ruppert averages, this functional result captures the temporal structure of the fluctuations and applies to non-smooth settings such as robust location estimation, including the geometric median.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。