用专家和普通演示数据提升强化学习训练速度
Skill-Enhanced Reinforcement Learning Acceleration from Heterogeneous Demonstrations
- 分两阶段:先从多种演示中提取技能先验知识
- 在多个基准上显著加速下游强化学习,尤其初期效果明显
- 适合数据稀缺场景,可提升技能学习效率
从演示中学习(LfD)是强化学习中的经典问题,旨在通过专家演示预训练智能体以加速后续学习。然而专家数据有限,制约了其效果。为此,我们提出新颖的两阶段方法——技能增强强化学习加速(SeRLA)。SeRLA 在离线先验学习阶段,引入技能级对抗性正-未标记(PU)学习模型,从专家演示和低成本通用演示中联合学习有用技能先验。在在线下游强化学习阶段,基于技能的软演员-评论家算法利用这些先验高效训练技能策略网络。此外,我们设计了一种简单的技能级数据增强技术,缓解数据稀疏问题,进一步提升先验学习与策略训练效果。在多个标准强化学习基准上的实验表明,SeRLA 在下游任务中实现了领先的强化学习加速性能,尤其在训练早期阶段优势显著。
原文摘要 · Abstract (English)
Learning from Demonstration (LfD) is a well-established problem in Reinforcement Learning (RL), which aims to facilitate rapid RL by leveraging expert demonstrations to pre-train the RL agent. However, the limited availability of expert demonstration data often hinders its ability to effectively aid downstream RL learning. To address this problem, we propose a novel two-stage method dubbed as Skill-enhanced Reinforcement Learning Acceleration (SeRLA). SeRLA introduces a skill-level adversarial Positive-Unlabeled (PU) learning model that extracts useful skill prior knowledge by learning from both expert demonstrations and general low-cost demonstrations in the offline prior learning stage. Building on this, it employs a skill-based soft actor-critic algorithm to leverage the acquired priors for efficient training of a skill policy network in the downstream online RL stage. In addition, we propose a simple skill-level data enhancement technique to mitigate data sparsity and further improve both skill prior learning and skill policy training. Experiments across multiple standard RL benchmarks demonstrate that SeRLA achieves state-of-the-art performance in accelerating reinforcement learning on downstream tasks, particularly in the early training phase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。