arXiv:2505.00787cs.LGcs.AI2025-05NeurIPS被引 4

提出最优行为基底,让多任务强化学习零样本快速找到最佳解。

Constructing an Optimal Behavior Basis for the Option Keyboard

  • 通过学习元策略动态组合基础策略,构建最优行为基底。
  • 减少所需基础策略数量,且在复杂任务中表现显著优于现有方法。
  • 可解决线性与部分非线性任务,适合多任务强化学习研究者。

多任务强化学习旨在以极少甚至无环境交互快速解决新任务。广义策略改进(GPI)通过组合一组基础策略生成新策略,使其至少不差于任一基础策略。在线性奖励场景下,可通过计算凸覆盖集(CCS)确保最优性,但该方法计算成本高,难以扩展至复杂领域。选项键盘(OK)改进了GPI,能生成至少不差、通常更优的策略,依赖于学习到的元策略动态组合基础策略。然而其性能高度依赖基础策略的选择。本文提出核心问题:是否存在一个最优基础策略集合——即最优行为基底——可实现任意线性任务的零样本最优解?我们提出一种新方法,高效构建此类基底,证明其显著减少达成最优所需的基底策略数量,并严格优于CCS,可最优求解特定非线性任务。我们在挑战性环境中进行实验,结果表明该方法优于当前最优技术,且随着任务复杂度提升,优势愈发明显。

原文摘要 · Abstract (English)

Multi-task reinforcement learning aims to quickly identify solutions for new tasks with minimal or no additional interaction with the environment. Generalized Policy Improvement (GPI) addresses this by combining a set of base policies to produce a new one that is at least as good -- though not necessarily optimal -- as any individual base policy. Optimality can be ensured, particularly in the linear-reward case, via techniques that compute a Convex Coverage Set (CCS). However, these are computationally expensive and do not scale to complex domains. The Option Keyboard (OK) improves upon GPI by producing policies that are at least as good -- and often better. It achieves this through a learned meta-policy that dynamically combines base policies. However, its performance critically depends on the choice of base policies. This raises a key question: is there an optimal set of base policies -- an optimal behavior basis -- that enables zero-shot identification of optimal solutions for any linear tasks? We solve this open problem by introducing a novel method that efficiently constructs such an optimal behavior basis. We show that it significantly reduces the number of base policies needed to ensure optimality in new tasks. We also prove that it is strictly more expressive than a CCS, enabling particular classes of non-linear tasks to be solved optimally. We empirically evaluate our technique in challenging domains and show that it outperforms state-of-the-art approaches, increasingly so as task complexity increases.

强化学习多任务策略组合最优基底

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。