arXiv:2602.15473cs.LG2026-02

用强化学习预测自适应学习率,提升优化器泛化能力

POP: Prior-Fitted First-Order Optimization Policies

  • 通过强化学习策略根据优化轨迹动态预测学习率
  • 在43个函数上超越传统梯度优化方法,无需任务调参
  • 基于百万合成问题训练,支持复杂场景的强泛化

基于梯度的优化器对自适应学习率机制的设计选择高度敏感。为解决这一问题,我们提出POP——一种元学习的强化学习策略,用于根据优化轨迹提供的上下文信息预测梯度下降的自适应学习率。该方法引入了新的强化学习奖励设计、一种面向分布内泛化的函数缩放策略,以及用于生成百万个合成优化问题的新先验。我们在包含43种不同复杂度优化函数的基准测试中评估POP,结果表明其显著优于现有梯度优化方法,且无需针对具体任务进行调参即可实现强大泛化能力。

原文摘要 · Abstract (English)

Gradient-based optimizers are highly sensitive to design choices in their adaptive learning rate mechanisms. To address this limitation, we introduce POP, a meta-learned Reinforcement Learning (RL) policy that predicts adaptive learning rates for gradient descent, conditioned on the contextual information provided in the optimization trajectory. Our method introduces a novel RL reward formulation, a new function-scaling strategy for in-distribution generalization, and a novel prior that is used to sample millions of synthetic optimization problems. We evaluate POP on an established benchmark including 43 optimization functions of various complexity, where it significantly outperforms gradient-based methods. Our evaluation demonstrates strong generalization capabilities without task-specific tuning.

优化算法强化学习元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。