arXiv:2410.00051cs.LGcs.AI2024-10NeurIPS被引 12

将一致性策略拓展至视觉强化学习,提升样本效率与训练稳定性。

Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization

  • 引入优先级近端经验正则化,稳定高维状态下的策略训练。
  • 在DeepMind Control Suite和Meta-world的21个任务中达新SOTA性能。
  • 首个将扩散/一致性模型应用于视觉强化学习的方法,适合高效视觉决策研究者。

高维状态空间下,视觉强化学习在利用与探索方面面临巨大挑战,导致样本效率低且训练不稳定。尽管一致性模型作为高效扩散模型已在基于状态的在线强化学习中验证有效,但其能否拓展至视觉强化学习仍是未解问题。本文研究非平稳分布及演员-评论家框架对一致性策略的影响,发现其在训练中不稳定,尤其在高维状态空间的视觉强化学习中。为此,我们提出基于样本熵正则化的策略训练稳定方法,并构建一致性策略与优先级近端经验正则化(CP3ER)框架,显著提升样本效率。CP3ER在DeepMind Control Suite和Meta-world的21个任务中取得新SOTA表现。据我们所知,这是首个将扩散/一致性模型应用于视觉强化学习的方法,展示了该类模型在视觉强化学习中的潜力。更多可视化结果见https://jzndd.github.io/CP3ER-Page/。

原文摘要 · Abstract (English)

With high-dimensional state spaces, visual reinforcement learning (RL) faces significant challenges in exploitation and exploration, resulting in low sample efficiency and training stability. As a time-efficient diffusion model, although consistency models have been validated in online state-based RL, it is still an open question whether it can be extended to visual RL. In this paper, we investigate the impact of non-stationary distribution and the actor-critic framework on consistency policy in online RL, and find that consistency policy was unstable during the training, especially in visual RL with the high-dimensional state space. To this end, we suggest sample-based entropy regularization to stabilize the policy training, and propose a consistency policy with prioritized proximal experience regularization (CP3ER) to improve sample efficiency. CP3ER achieves new state-of-the-art (SOTA) performance in 21 tasks across DeepMind control suite and Meta-world. To our knowledge, CP3ER is the first method to apply diffusion/consistency models to visual RL and demonstrates the potential of consistency models in visual RL. More visualization results are available at https://jzndd.github.io/CP3ER-Page/.

视觉RL一致性模型策略优化样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。