arXiv:2410.09728cs.LG2024-10NeurIPS被引 7

提出新元强化学习框架,可高效适配新任务并保证近最优性能。

Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum Comparator

  • 基于双层优化设计元策略,一次数据收集实现多步策略优化。
  • 理论证明算法在任务分布上的期望最优差距有上界,衡量泛化能力。
  • 实验验证理论正确性,性能优于现有基准方法。

元强化学习(Meta-RL)因其提升强化学习算法数据效率与泛化能力的潜力而受到关注。本文提出一种双层优化元强化学习框架(BO-MRL),用于学习任务特定策略适配的元先验,可在单次数据采集后实现多步策略优化。不同于现有分析,我们给出了任务分布上期望最优差距的上界,该指标衡量从元先验出发的策略适配与任务最优解之间的距离,从而量化模型对任务分布的泛化能力。我们通过实验证明了所推导上界的正确性,并展示了该算法在基准测试中的显著优越性。

原文摘要 · Abstract (English)

Meta-reinforcement learning (Meta-RL) has attracted attention due to its capability to enhance reinforcement learning (RL) algorithms, in terms of data efficiency and generalizability. In this paper, we develop a bilevel optimization framework for meta-RL (BO-MRL) to learn the meta-prior for task-specific policy adaptation, which implements multiple-step policy optimization on one-time data collection. Beyond existing meta-RL analyses, we provide upper bounds of the expected optimality gap over the task distribution. This metric measures the distance of the policy adaptation from the learned meta-prior to the task-specific optimum, and quantifies the model's generalizability to the task distribution. We empirically validate the correctness of the derived upper bounds and demonstrate the superior effectiveness of the proposed algorithm over benchmarks.

元强化学习双层优化策略适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。