用环境机制调节扩散模型,提升离线强化学习轨迹生成质量
Diffusion Modulation via Environment Mechanism Modeling for Planning
- 通过建模环境的转移动态与奖励函数,调制扩散训练过程
- 在多个离线RL任务上达到当前最优规划性能
- 适合需要高一致性轨迹生成的研究者与工程应用
扩散模型在离线强化学习中的轨迹生成方面展现出巨大潜力。然而,传统基于扩散的规划方法往往忽略了一个关键问题:强化学习中的轨迹生成需要确保状态转移的一致性,以保证在真实环境中具有连贯性。这一疏忽会导致生成轨迹与真实环境机制存在显著差异。为此,本文提出一种新型扩散规划方法——基于环境机制建模的扩散调制(DMEMM)。DMEMM通过引入强化学习环境的关键机制,特别是转移动态和奖励函数,对扩散模型训练进行调制。实验表明,该方法在多个离线强化学习任务中均实现了当前最优的规划性能。
原文摘要 · Abstract (English)
Diffusion models have shown promising capabilities in trajectory generation for planning in offline reinforcement learning (RL). However, conventional diffusion-based planning methods often fail to account for the fact that generating trajectories in RL requires unique consistency between transitions to ensure coherence in real environments. This oversight can result in considerable discrepancies between the generated trajectories and the underlying mechanisms of a real environment. To address this problem, we propose a novel diffusion-based planning method, termed as Diffusion Modulation via Environment Mechanism Modeling (DMEMM). DMEMM modulates diffusion model training by incorporating key RL environment mechanisms, particularly transition dynamics and reward functions. Experimental results demonstrate that DMEMM achieves state-of-the-art performance for planning with offline reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。