首个针对三重优化的随机梯度方法,理论完备且适应各种误差。
A stochastic gradient method for trilevel optimization
- 提出首个无约束三重优化的随机梯度下降算法
- 在中层、底层求解不精确及梯度噪声下仍保证收敛
- 适用于超参数对抗调优等复杂机器学习任务
随着双层优化领域的成功,更具挑战性的三重优化问题也逐渐受到关注。本文首次提出用于求解无约束三重优化问题的随机梯度下降方法,并建立了涵盖各类不精确性的收敛理论:包括中层与底层问题的近似解、三重伴随公式计算误差,以及梯度、海森矩阵、雅可比矩阵和三阶张量的噪声估计。通过合成三重问题与超参数对抗调优的数值实验,验证了该方法的有效性。
原文摘要 · Abstract (English)
With the success that the field of bilevel optimization has seen in recent years, similar methodologies have started being applied to solving more difficult applications that arise in trilevel optimization. At the helm of these applications are new machine learning formulations that have been proposed in the trilevel context and, as a result, efficient and theoretically sound stochastic methods are required. In this work, we propose the first-ever stochastic gradient descent method for solving unconstrained trilevel optimization problems and provide a convergence theory that covers all forms of inexactness of the trilevel adjoint gradient, such as the inexact solutions of the middle-level and lower-level problems, inexact computation of the trilevel adjoint formula, and noisy estimates of the gradients, Hessians, Jacobians, and tensors of third-order derivatives involved. We also demonstrate the promise of our approach by providing numerical results on both synthetic trilevel problems and trilevel formulations for hyperparameter adversarial tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。