揭示标准SGD在多指标模型中的根本局限,打破传统统计查询框架的误导。
Limitations of SGD for Multi-Index Models Beyond Statistical Queries
- 提出非统计查询新框架,直接分析原生SGD的性能边界。
- 证明在低维投影模型中,标准SGD存在无法突破的泛化误差下界。
- 适用于深层网络等复杂架构,为理解优化困境提供新视角。
理解梯度方法、尤其是随机梯度下降(SGD)的局限性是学习理论的核心挑战。常用工具统计查询(SQ)框架通过数据噪声交互研究算法性能上限,但其与SGD的理论联系薄弱:现有结果常依赖对抗性或特殊结构的梯度噪声,这些噪声不反映标准SGD的实际行为,甚至可能导致错误预测。此外,许多关于复杂问题中SGD的分析依赖非平凡的算法修改,如将轨迹限制在球面或使用极小学习率。为克服这些缺陷,本文构建了一个新的非统计查询框架,用于研究单指标和多指标模型(即目标函数依赖输入的低维投影)中标准原生SGD的性能极限。该框架适用于广泛设置与架构,包括潜在深层神经网络。
原文摘要 · Abstract (English)
Understanding the limitations of gradient methods, and stochastic gradient descent (SGD) in particular, is a central challenge in learning theory. To that end, a commonly used tool is the Statistical Queries (SQ) framework, which studies performance limits of algorithms based on noisy interaction with the data. However, it is known that the formal connection between the SQ framework and SGD is tenuous: Existing results typically rely on adversarial or specially-structured gradient noise that does not reflect the noise in standard SGD, and (as we point out here) can sometimes lead to incorrect predictions. Moreover, many analyses of SGD for challenging problems rely on non-trivial algorithmic modifications, such as restricting the SGD trajectory to the sphere or using very small learning rates. To address these shortcomings, we develop a new, non-SQ framework to study the limitations of standard vanilla SGD, for single-index and multi-index models (namely, when the target function depends on a low-dimensional projection of the inputs). Our results apply to a broad class of settings and architectures, including (potentially deep) neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。