通过引入曲率信息提升双层优化超梯度计算效率。
Efficient Curvature-Aware Hypergradient Approximation for Bilevel Optimization
- 利用曲率信息改进超梯度近似,结合不精确牛顿法
- 理论证明在确定性场景下计算复杂度更低
- 适合需要高效超参数优化的机器学习任务
双层优化是超参数优化和元学习等机器学习问题的强大工具。估计超梯度(又称隐式梯度)对于发展基于梯度的双层优化方法至关重要。本文提出一种计算高效的技巧,将曲率信息融入超梯度近似,并基于此构建新型算法框架。该框架在确定性和随机场景下均提供收敛速率保证,尤其在确定性设置中相比主流梯度方法实现了更优的计算复杂度。这一复杂度提升源于对超梯度结构的精细利用及不精确牛顿法的应用。除理论加速外,数值实验也验证了引入曲率信息带来的显著实际性能优势。
原文摘要 · Abstract (English)
Bilevel optimization is a powerful tool for many machine learning problems, such as hyperparameter optimization and meta-learning. Estimating hypergradients (also known as implicit gradients) is crucial for developing gradient-based methods for bilevel optimization. In this work, we propose a computationally efficient technique for incorporating curvature information into the approximation of hypergradients and present a novel algorithmic framework based on the resulting enhanced hypergradient computation. We provide convergence rate guarantees for the proposed framework in both deterministic and stochastic scenarios, particularly showing improved computational complexity over popular gradient-based methods in the deterministic setting. This improvement in complexity arises from a careful exploitation of the hypergradient structure and the inexact Newton method. In addition to the theoretical speedup, numerical experiments demonstrate the significant practical performance benefits of incorporating curvature information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。