arXiv:2412.06436math.OCcs.LG2024-12被引 4

提出自适应不精确方法,高效优化双层学习中的可学习算子。

An Adaptively Inexact Method for Bilevel Learning Using Primal-Dual Style Differentiation

  • 用猪背法计算梯度与下层问题解,容忍数值误差。
  • 推导后验误差界,指导下层求解精度与步长设置。
  • 适合需学习正则化器或神经网络结构的研究者。

本文研究用于学习线性算子的双层学习框架。在此框架中,可学习参数通过依赖于凸优化问题(即下层问题)最小化解的损失函数进行优化。我们采用一种称为'piggyback'的迭代算法来计算损失函数及其梯度。由于下层问题需数值求解,损失函数及梯度只能近似计算。为此,我们推导了后验误差界,用于估计超梯度计算的准确性,并据此指导下层问题的容差以及piggyback算法的设置。为进一步提升上层优化效率,我们还提出了自适应步长选择方法。通过若干学习正则化器的问题(如输入凸神经网络训练)验证了所提方法的有效性。

原文摘要 · Abstract (English)

We consider a bilevel learning framework for learning linear operators. In this framework, the learnable parameters are optimized via a loss function that also depends on the minimizer of a convex optimization problem (denoted lower-level problem). We utilize an iterative algorithm called `piggyback' to compute the gradient of the loss and minimizer of the lower-level problem. Given that the lower-level problem is solved numerically, the loss function and thus its gradient can only be computed inexactly. To estimate the accuracy of the computed hypergradient, we derive an a-posteriori error bound, which provides guides for setting the tolerance for the lower-level problem, as well as the piggyback algorithm. To efficiently solve the upper-level optimization, we also propose an adaptive method for choosing a suitable step-size. To illustrate the proposed method, we consider a few learned regularizer problems, such as training an input-convex neural network.

双层学习优化算法神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。