用模型生成更合适的任务目标,提升离线强化学习效果
MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning
- 基于学习的动力学模型生成新目标,增强数据多样性
- 在状态和视觉迷宫任务中性能显著超越已有方法
- 适合做离线目标条件强化学习的研究者和开发者
近期提出的面向离线目标条件强化学习的先进算法——目标条件加权监督学习(GCWSL),通过优化目标条件强化学习目标的下界,在多种目标达成任务中表现出色。然而,现有方法缺乏轨迹拼接能力。为解决此问题,已有研究尝试通过目标数据增强来改进,但难以有效采样适合的增强目标。本文提出基于模型的目标数据增强(MGDA)方法,依据目标多样性、动作最优性和目标可达性三大原则,利用学习到的动力学模型生成更优的增强目标。MGDA引入局部利普希茨连续性假设,缓解误差累积影响。实验表明,该方法显著提升了GCWSL在状态与视觉迷宫数据集上的表现,优于以往增强技术,显著改善了拼接能力。
原文摘要 · Abstract (English)
Recently, a state-of-the-art family of algorithms, known as Goal-Conditioned Weighted Supervised Learning (GCWSL) methods, has been introduced to tackle challenges in offline goal-conditioned reinforcement learning (RL). GCWSL optimizes a lower bound of the goal-conditioned RL objective and has demonstrated outstanding performance across diverse goal-reaching tasks, providing a simple, effective, and stable solution. However, prior research has identified a critical limitation of GCWSL: the lack of trajectory stitching capabilities. To address this, goal data augmentation strategies have been proposed to enhance these methods. Nevertheless, existing techniques often struggle to sample suitable augmented goals for GCWSL effectively. In this paper, we establish unified principles for goal data augmentation, focusing on goal diversity, action optimality, and goal reachability. Based on these principles, we propose a Model-based Goal Data Augmentation (MGDA) approach, which leverages a learned dynamics model to sample more suitable augmented goals. MGDA uniquely incorporates the local Lipschitz continuity assumption within the learned model to mitigate the impact of compounding errors. Empirical results show that MGDA significantly enhances the performance of GCWSL methods on both state-based and vision-based maze datasets, surpassing previous goal data augmentation techniques in improving stitching capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。