arXiv:2603.01951cs.LGmath.OC2026-03

首次在流式线性预测中实现动量加速,突破传统方法瓶颈。

Accelerating Single-Pass SGD for Generalized Linear Prediction

  • 用数据相关近端法引入动量,实现双重加速
  • 优化误差更优,统计误差达最优,模型偏差可控
  • 适合追求高效在线学习的科研与工程人员

我们研究流式设置下的广义线性预测,每轮仅使用一个新数据点进行梯度更新。尽管动量在确定性优化中已被广泛证实有效,但在单遍非二次随机优化中能否加速仍是一个基本开放问题。本文提出首个通过新颖的数据依赖近端方法成功引入动量的算法,实现双动量加速。所推导的过量风险界分解为三部分:改进的优化误差、极小化最优的统计误差,以及高阶模型误设误差。证明通过细粒度的内循环平稳性分析处理误设问题,同时通过两阶段外循环分析定位统计误差。最终解决Jain等[2018a]提出的开放问题,表明在流式设定下,动量加速比方差缩减对广义线性预测更有效。

原文摘要 · Abstract (English)

We study generalized linear prediction under a streaming setting, where each iteration uses only one fresh data point for a gradient-level update. While momentum is well-established in deterministic optimization, a fundamental open question is whether it can accelerate such single-pass non-quadratic stochastic optimization. We propose the first algorithm that successfully incorporates momentum via a novel data-dependent proximal method, achieving dual-momentum acceleration. Our derived excess risk bound decomposes into three components: an improved optimization error, a minimax optimal statistical error, and a higher-order model-misspecification error. The proof handles mis-specification via a fine-grained stationary analysis of inner updates, while localizing statistical error through a two-phase outer-loop analysis. As a result, we resolve the open problem posed by Jain et al. [2018a] and demonstrate that momentum acceleration is more effective than variance reduction for generalized linear prediction in the streaming setting.

在线学习动量加速统计优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。