arXiv:2507.11228cs.LGmath.OC2025-07被引 4

研究大步长梯度下降在球面数据上的收敛性,发现高维下仍会震荡。

Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?

  • 将数据限制在相同幅值的球面上,分析梯度下降行为
  • 一维情况下可全局收敛,但高维中仍会出现循环震荡
  • 为理解真实数据中的震荡现象提供理论基础

对逻辑回归的梯度下降(GD)存在许多有趣性质。当数据线性可分时,无论步长多大,迭代点方向都会收敛到最大间隔分离超平面。但在不可分情况下,已有研究表明,即使步长低于稳定性阈值 $2/λ$($λ$ 为解处海森矩阵的最大特征值),GD 仍可能出现循环行为。本文探讨:若将数据限制在等幅球面上,是否足以保证在任意低于稳定性阈值的步长下实现全局收敛。我们证明在一维情形下成立,但在高维空间中,循环行为仍可能发生。希望激发更多研究,以量化真实数据中此类循环的普遍性,并寻找确保大步长全局收敛的充分条件。

原文摘要 · Abstract (English)

Gradient descent (GD) on logistic regression has many fascinating properties. When the dataset is linearly separable, it is known that the iterates converge in direction to the maximum-margin separator regardless of how large the step size is. In the non-separable case, however, it has been shown that GD can exhibit a cycling behaviour even when the step sizes is still below the stability threshold $2/λ$, where $λ$ is the largest eigenvalue of the Hessian at the solution. This short paper explores whether restricting the data to have equal magnitude is a sufficient condition for global convergence, under any step size below the stability threshold. We prove that this is true in a one dimensional space, but in higher dimensions cycling behaviour can still occur. We hope to inspire further studies on quantifying how common these cycles are in realistic datasets, as well as finding sufficient conditions to guarantee global convergence with large step sizes.

优化理论梯度下降逻辑回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。