arXiv:2412.06207cs.LGcs.AI2024-12被引 1

用专家和普通演示数据提升强化学习训练速度

Skill-Enhanced Reinforcement Learning Acceleration from Heterogeneous Demonstrations

  • 分两阶段:先从多种演示中提取技能先验知识
  • 在多个基准上显著加速下游强化学习,尤其初期效果明显
  • 适合数据稀缺场景,可提升技能学习效率

从演示中学习(LfD)是强化学习中的经典问题,旨在通过专家演示预训练智能体以加速后续学习。然而专家数据有限,制约了其效果。为此,我们提出新颖的两阶段方法——技能增强强化学习加速(SeRLA)。SeRLA 在离线先验学习阶段,引入技能级对抗性正-未标记(PU)学习模型,从专家演示和低成本通用演示中联合学习有用技能先验。在在线下游强化学习阶段,基于技能的软演员-评论家算法利用这些先验高效训练技能策略网络。此外,我们设计了一种简单的技能级数据增强技术,缓解数据稀疏问题,进一步提升先验学习与策略训练效果。在多个标准强化学习基准上的实验表明,SeRLA 在下游任务中实现了领先的强化学习加速性能,尤其在训练早期阶段优势显著。

原文摘要 · Abstract (English)

Learning from Demonstration (LfD) is a well-established problem in Reinforcement Learning (RL), which aims to facilitate rapid RL by leveraging expert demonstrations to pre-train the RL agent. However, the limited availability of expert demonstration data often hinders its ability to effectively aid downstream RL learning. To address this problem, we propose a novel two-stage method dubbed as Skill-enhanced Reinforcement Learning Acceleration (SeRLA). SeRLA introduces a skill-level adversarial Positive-Unlabeled (PU) learning model that extracts useful skill prior knowledge by learning from both expert demonstrations and general low-cost demonstrations in the offline prior learning stage. Building on this, it employs a skill-based soft actor-critic algorithm to leverage the acquired priors for efficient training of a skill policy network in the downstream online RL stage. In addition, we propose a simple skill-level data enhancement technique to mitigate data sparsity and further improve both skill prior learning and skill policy training. Experiments across multiple standard RL benchmarks demonstrate that SeRLA achieves state-of-the-art performance in accelerating reinforcement learning on downstream tasks, particularly in the early training phase.

强化学习技能学习演示学习数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。