arXiv:2505.17304cs.LGmath.OC2025-05

揭示梯度下降在过参数模型中的隐式正则化机制

Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models

  • 通过微小扰动设计新型优化算法IPGD
  • 理论证明其在矩阵感知任务中收敛至低维解空间
  • 适合研究优化动态与过参数模型泛化性的读者

隐式正则化指局部搜索算法倾向于收敛到低维结构解,即使未显式施加此类约束。尽管普遍存在,其内在机制在过参数化设置下仍不清晰。本文分析梯度下降动力学,识别出三种使其收敛至二阶平稳点并保持在隐式低维区域的条件:(i) 合适初始化,(ii) 高效逃离鞍点,(iii) 持续靠近该区域。我们证明这些条件可通过无穷小扰动和小偏差率实现。基于此,提出无穷小扰动梯度下降(IPGD),在温和假设下满足上述条件。理论保证了其在过参数化矩阵感知任务中的表现,并提供了广泛适用性的实证证据。

原文摘要 · Abstract (English)

Implicit regularization refers to the tendency of local search algorithms to converge to low-dimensional solutions, even when such structures are not explicitly enforced. Despite its ubiquity, the mechanism underlying this behavior remains poorly understood, particularly in over-parameterized settings. We analyze gradient descent dynamics and identify three conditions under which it converges to second-order stationary points within an implicit low-dimensional region: (i) suitable initialization, (ii) efficient escape from saddle points, and (iii) sustained proximity to the region. We show that these can be achieved through infinitesimal perturbations and a small deviation rate. Building on this, we introduce Infinitesimally Perturbed Gradient Descent (IPGD), which satisfies these conditions under mild assumptions. We provide theoretical guarantees for IPGD in over-parameterized matrix sensing and empirical evidence of its broader applicability.

优化算法隐式正则化过参数化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。