用扩散模型融合异构数据,提升离线强化学习泛化能力
Off-dynamics Conditional Diffusion Planners
- 用条件扩散模型联合建模异源数据与目标数据分布
- 在多种环境上性能超越多个强基线方法
- 可调节上下文实现动态迁移与鲁棒性增强
离线强化学习通过利用已有数据集避免交互式数据采集,但其效果依赖于数据的数量与质量。本文探索使用更易获取的异动力学数据集来缓解离线强化学习中的数据稀缺问题。提出一种基于条件扩散概率模型(Conditional DPMs)的新方法,用于学习大规模异动力学数据集与有限目标数据集的联合分布。为使模型捕捉底层动力学结构,引入两种条件上下文:(1) 连续动力学得分允许两类数据轨迹部分重叠,提供更丰富信息;(2) 反向动力学上下文引导模型生成符合目标环境动态约束的轨迹。实验表明,该方法显著优于多个强基线。消融实验进一步揭示各动力学上下文的关键作用。此外,通过调整上下文,模型可实现源与目标动态间的插值,对环境微小变化更具鲁棒性。
原文摘要 · Abstract (English)
Offline Reinforcement Learning (RL) offers an attractive alternative to interactive data acquisition by leveraging pre-existing datasets. However, its effectiveness hinges on the quantity and quality of the data samples. This work explores the use of more readily available, albeit off-dynamics datasets, to address the challenge of data scarcity in Offline RL. We propose a novel approach using conditional Diffusion Probabilistic Models (DPMs) to learn the joint distribution of the large-scale off-dynamics dataset and the limited target dataset. To enable the model to capture the underlying dynamics structure, we introduce two contexts for the conditional model: (1) a continuous dynamics score allows for partial overlap between trajectories from both datasets, providing the model with richer information; (2) an inverse-dynamics context guides the model to generate trajectories that adhere to the target environment's dynamic constraints. Empirical results demonstrate that our method significantly outperforms several strong baselines. Ablation studies further reveal the critical role of each dynamics context. Additionally, our model demonstrates that by modifying the context, we can interpolate between source and target dynamics, making it more robust to subtle shifts in the environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。