arXiv:2601.14647math.OCcs.LG2026-01

用方差缩减的随机信任域法加速非凸优化,无需函数值即可收敛。

TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction

  • 基于随机梯度与方差缩减技术自适应调整信任域半径。
  • 在温和假设下收敛到一阶驻点,迭代与样本复杂度匹配最优水平。
  • 适合需要高效二阶信息的机器学习任务,尤其对批量大小敏感。

我们提出一种用于无约束非凸优化的随机信任域方法,引入方差缩减梯度(SVRG)以加速收敛。与经典信任域方法不同,该算法仅依赖随机梯度信息,无需函数值评估。信任域半径根据半径控制参数和随机梯度估计自适应调整。在温和假设下,算法期望收敛至一阶驻点。此外,该方法达到与基于SVRG的一阶方法相当的迭代与样本复杂度,同时支持随机且可能依赖梯度的二阶信息。大量数值实验表明,引入SVRG可加速收敛,使用信任域及海森矩阵信息进一步提升性能。我们还分析了批大小与内循环长度对效率的影响,结果显示该方法在多个机器学习任务中优于SGD与Adam。

原文摘要 · Abstract (English)

We propose a stochastic trust-region method for unconstrained nonconvex optimization that incorporates stochastic variance-reduced gradients (SVRG) to accelerate convergence. Unlike classical trust-region methods, the proposed algorithm relies solely on stochastic gradient information and does not require function value evaluations. The trust-region radius is adaptively adjusted based on a radius-control parameter and the stochastic gradient estimate. Under mild assumptions, we establish that the algorithm converges in expectation to a first-order stationary point. Moreover, the method achieves iteration and sample complexity bounds that match those of SVRG-based first-order methods, while allowing stochastic and potentially gradient-dependent second-order information. Extensive numerical experiments demonstrate that incorporating SVRG accelerates convergence, and that the use of trust-region methods and Hessian information further improves performance. We also highlight the impact of batch size and inner-loop length on efficiency, and show that the proposed method outperforms SGD and Adam on several machine learning tasks.

优化算法信任域方差缩减非凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。