arXiv:2410.11312cs.LGcs.AI2024-10

提出一种新型梯度法,高效解决多层级优化难题

Towards Differentiable Multilevel Optimization: A Gradient-Based Approach

  • 通过分层分解梯度并改进传播机制,突破嵌套结构计算瓶颈
  • 在多层级场景下显著降低复杂度,准确率与收敛速度均提升
  • 首个兼具理论保障与优异实证表现的隐式微分通用算法

多层级优化因在超参数调优和持续学习等任务中的潜力而重获关注。然而,现有方法难以高效处理其固有的嵌套结构。本文提出一种新型基于梯度的多层级优化方法,通过层次化分解完整梯度并采用先进传播技术,克服上述局限。该方法可扩展至 n 级场景,在显著降低计算复杂度的同时,提升解的准确率与收敛速度。通过多个基准测试的数值实验对比,结果表明解的准确率有明显提升。据我们所知,这是首个同时具备理论保证与优越实证性能的通用隐式微分算法。

原文摘要 · Abstract (English)

Multilevel optimization has gained renewed interest in machine learning due to its promise in applications such as hyperparameter tuning and continual learning. However, existing methods struggle with the inherent difficulty of efficiently handling the nested structure. This paper introduces a novel gradient-based approach for multilevel optimization that overcomes these limitations by leveraging a hierarchically structured decomposition of the full gradient and employing advanced propagation techniques. Extending to n-level scenarios, our method significantly reduces computational complexity while improving both solution accuracy and convergence speed. We demonstrate the effectiveness of our approach through numerical experiments, comparing it with existing methods across several benchmarks. The results show a notable improvement in solution accuracy. To the best of our knowledge, this is one of the first algorithms to provide a general version of implicit differentiation with both theoretical guarantees and superior empirical performance.

多层级优化梯度法隐式微分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。