提出一种高效随机双层学习方法,解决图像去噪与去模糊的超参数优化问题。
Bilevel Learning with Inexact Stochastic Gradients
- 采用数据采样生成的不精确随机梯度,避免固定下层迭代次数
- 在弱假设下证明收敛性,实测速度提升且泛化能力更强
- 适合大规模图像重建任务,尤其对计算效率要求高的场景
双层学习在机器学习、反问题和成像应用中日益重要,涵盖超参数优化、自适应正则化器学习以及前向算子优化。由于这些问题规模大,已有研究发展出不精确且计算高效的算法。现有自适应方法多基于确定性框架,而随机方法通常依赖不切实际的方差假设,强制固定下层迭代次数并需大量调参。本文针对下层问题强凸、上层目标为非凸函数之和的情况,考虑上层数据采样带来的随机性,导致不精确的随机超梯度。我们建立了其与非凸随机优化前沿理论的联系,并在温和假设下证明了不精确随机双层优化的收敛性。实验结果表明,在图像去噪与去模糊等成像任务中,该方法相较自适应确定性双层方法显著提速并提升泛化性能。
原文摘要 · Abstract (English)
Bilevel learning has gained prominence in machine learning, inverse problems, and imaging applications, including hyperparameter optimization, learning data-adaptive regularizers, and optimizing forward operators. The large-scale nature of these problems has led to the development of inexact and computationally efficient methods. Existing adaptive methods predominantly rely on deterministic formulations, while stochastic approaches often adopt a doubly-stochastic framework with impractical variance assumptions, enforces a fixed number of lower-level iterations, and requires extensive tuning. In this work, we focus on bilevel learning with strongly convex lower-level problems and a nonconvex sum-of-functions in the upper-level. Stochasticity arises from data sampling in the upper-level which leads to inexact stochastic hypergradients. We establish their connection to state-of-the-art stochastic optimization theory for nonconvex objectives. Furthermore, we prove the convergence of inexact stochastic bilevel optimization under mild assumptions. Our empirical results highlight significant speed-ups and improved generalization in imaging tasks such as image denoising and deblurring in comparison with adaptive deterministic bilevel methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。