arXiv:2412.20519cs.LGcs.AI2024-12被引 5

用目标引导的生成方法,提升低质量离线数据集的训练效果

Goal-Conditioned Data Augmentation for Offline Reinforcement Learning

  • 基于扩散模型,根据回报导向选择高价值目标生成新数据
  • 在D4RL和交通信号控制任务上,显著提升策略性能
  • 适合数据稀缺或示范质量差的离线强化学习场景

离线强化学习(Offline RL)允许从预先收集的数据集中学习策略,无需与环境直接交互。然而,受限于离线数据集的质量,其在次优数据集上通常难以学习到优质策略。为解决缺乏最优示范的问题,本文提出一种新的目标条件化数据增强方法——目标引导数据增强(GODA),该方法基于生成建模技术,引入回报导向的目标条件及多种选择机制。具体而言,通过可控缩放技术,在数据采样过程中提供更强的回报引导。GODA在学习原始数据集的完整分布的同时,生成具有选择性更高回报目标的新数据,从而最大化有限最优示范的利用效率。此外,提出一种自适应门控条件处理方法,以应对噪声输入与条件,增强目标导向的指导能力。在D4RL基准和真实世界交通信号控制(TSC)任务上的实验表明,GODA在提升数据质量方面表现优异,并在多种离线强化学习算法中优于当前最优数据增强方法。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) enables policy learning from pre-collected offline datasets, relaxing the need to interact directly with the environment. However, limited by the quality of offline datasets, it generally fails to learn well-qualified policies in suboptimal datasets. To address datasets with insufficient optimal demonstrations, we introduce Goal-cOnditioned Data Augmentation (GODA), a novel goal-conditioned diffusion-based method for augmenting samples with higher quality. Leveraging recent advancements in generative modelling, GODA incorporates a novel return-oriented goal condition with various selection mechanisms. Specifically, we introduce a controllable scaling technique to provide enhanced return-based guidance during data sampling. GODA learns a comprehensive distribution representation of the original offline datasets while generating new data with selectively higher-return goals, thereby maximizing the utility of limited optimal demonstrations. Furthermore, we propose a novel adaptive gated conditioning method for processing noisy inputs and conditions, enhancing the capture of goal-oriented guidance. We conduct experiments on the D4RL benchmark and real-world challenges, specifically traffic signal control (TSC) tasks, to demonstrate GODA's effectiveness in enhancing data quality and superior performance compared to state-of-the-art data augmentation methods across various offline RL algorithms.

离线强化学习数据增强生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。