arXiv:2506.10167cs.LGcs.SY2025-06被引 1

用最优与最差策略的混合提升强化学习采样效率

Wasserstein Barycenter Soft Actor-Critic

  • 结合悲观与乐观策略,通过Wasserstein均值生成探索策略
  • 在MuJoCo任务上显著提升样本效率,优于现有算法
  • 适合稀疏奖励场景的连续控制强化学习研究

深度离线策略演员-评论家算法已成为连续控制领域强化学习的主流框架。然而,大多数算法在稀疏奖励环境中仍存在样本效率低的问题。本文提出一种有原则的定向探索策略——Wasserstein Barycenter Soft Actor-Critic (WBSAC) 算法,该算法利用悲观演员进行时序差分学习,乐观演员促进探索,并通过悲观与乐观策略的Wasserstein均值生成探索策略,动态调整探索强度。在MuJoCo连续控制任务上的实验表明,WBSAC相较于当前最优的离线策略演员-评论家算法具有更高的样本效率。

原文摘要 · Abstract (English)

Deep off-policy actor-critic algorithms have emerged as the leading framework for reinforcement learning in continuous control domains. However, most of these algorithms suffer from poor sample efficiency, especially in environments with sparse rewards. In this paper, we take a step towards addressing this issue by providing a principled directed exploration strategy. We propose Wasserstein Barycenter Soft Actor-Critic (WBSAC) algorithm, which benefits from a pessimistic actor for temporal difference learning and an optimistic actor to promote exploration. This is achieved by using the Wasserstein barycenter of the pessimistic and optimistic policies as the exploration policy and adjusting the degree of exploration throughout the learning process. We compare WBSAC with state-of-the-art off-policy actor-critic algorithms and show that WBSAC is more sample-efficient on MuJoCo continuous control tasks.

强化学习策略优化采样效率连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。