arXiv:2510.17503cs.LGmath.OC2025-10被引 2

提出带动量的随机DC优化方法,小批量下也能保证收敛。

Stochastic Difference-of-Convex Optimization with Momentum

  • 引入动量机制,突破小批量收敛瓶颈
  • 理论证明动量对任意批大小都必要且充分
  • 兼具理论保障与实际性能,适合小批量学习场景

随机差分凸(DC)优化广泛应用于机器学习,但小批量下的收敛性仍不明确。现有方法通常需要大批次或强噪声假设,限制了实际应用。本文证明,在标准光滑性和凹部有界方差假设下,动量可使算法在任意批次大小下实现收敛。若无动量,无论步长如何,收敛可能失败,凸显其必要性。所提动量算法具备可证明的收敛性,并展现出强劲的实证表现。

原文摘要 · Abstract (English)

Stochastic difference-of-convex (DC) optimization is prevalent in numerous machine learning applications, yet its convergence properties under small batch sizes remain poorly understood. Existing methods typically require large batches or strong noise assumptions, which limit their practical use. In this work, we show that momentum enables convergence under standard smoothness and bounded variance assumptions (of the concave part) for any batch size. We prove that without momentum, convergence may fail regardless of stepsize, highlighting its necessity. Our momentum-based algorithm achieves provable convergence and demonstrates strong empirical performance.

优化算法随机优化动量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。