arXiv:2503.20176cs.LGcs.AI2025-03被引 4

用离散扩散技能提升离线强化学习的长时序任务表现

Offline Reinforcement Learning with Discrete Diffusion Skills

  • 构建基于变压器编码器与扩散解码器的离散技能空间
  • 在AntMaze-v2上比现有方法提升至少12%
  • 适合需要可解释性与稳定训练的复杂任务研究

技能作为时间抽象被引入离线强化学习,用于应对复杂、长时序任务,促进行为一致性并实现有意义的探索。尽管当前离线RL中的技能多在连续潜在空间建模,离散技能空间的潜力仍待挖掘。本文提出一种紧凑的离散技能空间,依托先进的基于Transformer的编码器和基于扩散的解码器。结合通过离线RL技术训练的高层策略,该方法建立了一个层次化强化学习框架,其中训练好的扩散解码器起关键作用。实验表明,所提算法离散扩散技能(DDS)是一种强大的离线强化学习方法,在运动与厨房任务上表现优异,并在长时序任务中显著领先,相较于现有方法在AntMaze-v2基准上至少提升12%。此外,相比以往基于技能的方法,DDS具备更优的可解释性、训练稳定性与在线探索能力。

原文摘要 · Abstract (English)

Skills have been introduced to offline reinforcement learning (RL) as temporal abstractions to tackle complex, long-horizon tasks, promoting consistent behavior and enabling meaningful exploration. While skills in offline RL are predominantly modeled within a continuous latent space, the potential of discrete skill spaces remains largely underexplored. In this paper, we propose a compact discrete skill space for offline RL tasks supported by state-of-the-art transformer-based encoder and diffusion-based decoder. Coupled with a high-level policy trained via offline RL techniques, our method establishes a hierarchical RL framework where the trained diffusion decoder plays a pivotal role. Empirical evaluations show that the proposed algorithm, Discrete Diffusion Skill (DDS), is a powerful offline RL method. DDS performs competitively on Locomotion and Kitchen tasks and excels on long-horizon tasks, achieving at least a 12 percent improvement on AntMaze-v2 benchmarks compared to existing offline RL approaches. Furthermore, DDS offers improved interpretability, training stability, and online exploration compared to previous skill-based methods.

离线强化学习扩散模型技能学习长时序任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。