arXiv:2410.14069cs.LG2024-10NeurIPS被引 8

用最优传输重构离线强化学习,从数据中拼接最优行为。

Rethinking Optimal Transport in Offline Reinforcement Learning

  • 将离线强化学习建模为最优传输问题,智能体从数据中选择最优动作。
  • 在D4RL连续控制任务上表现优于现有方法,提升策略效率。
  • 适合研究离线强化学习与多专家数据融合的学者使用。

我们提出一种基于最优传输的离线强化学习新算法。在离线强化学习中,数据通常由多个专家提供,其中部分专家行为可能次优。为提取高效策略,需从数据集中「拼接」最优行为。为此,我们将离线强化学习重新视为一个最优传输问题,并设计了一种算法,旨在找到一个策略,将状态映射到每个状态下最佳专家动作的部分分布。我们在D4RL基准的连续控制任务上评估了该算法,结果表明其性能优于现有方法。

原文摘要 · Abstract (English)

We propose a novel algorithm for offline reinforcement learning using optimal transport. Typically, in offline reinforcement learning, the data is provided by various experts and some of them can be sub-optimal. To extract an efficient policy, it is necessary to \emph{stitch} the best behaviors from the dataset. To address this problem, we rethink offline reinforcement learning as an optimal transportation problem. And based on this, we present an algorithm that aims to find a policy that maps states to a \emph{partial} distribution of the best expert actions for each given state. We evaluate the performance of our algorithm on continuous control problems from the D4RL suite and demonstrate improvements over existing methods.

强化学习最优传输离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。