arXiv:2502.12631cs.LGcs.AI2025-02ICML被引 5

用最优传输理论让扩散策略更稳定地结合强化学习。

Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport

  • 用最优传输理论将扩散策略与强化学习融合,提升鲁棒性。
  • 在三个仿真任务中表现优于现有方法,尤其在稀疏奖励环境下。
  • 适合需要高精度控制和长期规划的复杂任务研究者使用。

扩散策略在从示范中学习复杂行为方面展现出潜力,尤其适用于需要精确控制和长期规划的任务。然而,在遭遇分布偏移时,其鲁棒性较差。本文通过在线环境交互改进基于扩散的模仿学习模型。提出OTPR(基于最优传输的扩散策略强化学习微调),利用最优传输理论将扩散策略与强化学习结合。该方法以Q函数为传输代价,将策略视为最优传输映射,实现高效稳定的微调。此外,引入掩码最优传输以专家关键点引导状态-动作匹配,并采用基于兼容性的重采样策略增强训练稳定性。在三个仿真任务上的实验表明,OTPR在性能和鲁棒性上均优于现有方法,尤其在复杂且稀疏奖励环境中表现突出。总体而言,OTPR提供了一种有效框架,实现模仿学习与强化学习的融合,达成多功能且可靠的策略学习。代码将发布于https://github.com/Sunmmyy/OTPR.git。

原文摘要 · Abstract (English)

Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and long-term planning. However, they face challenges in robustness when encountering distribution shifts. This paper explores improving diffusion-based imitation learning models through online interactions with the environment. We propose OTPR (Optimal Transport-guided score-based diffusion Policy for Reinforcement learning fine-tuning), a novel method that integrates diffusion policies with RL using optimal transport theory. OTPR leverages the Q-function as a transport cost and views the policy as an optimal transport map, enabling efficient and stable fine-tuning. Moreover, we introduce masked optimal transport to guide state-action matching using expert keypoints and a compatibility-based resampling strategy to enhance training stability. Experiments on three simulation tasks demonstrate OTPR's superior performance and robustness compared to existing methods, especially in complex and sparse-reward environments. In sum, OTPR provides an effective framework for combining IL and RL, achieving versatile and reliable policy learning. The code will be released at https://github.com/Sunmmyy/OTPR.git.

扩散模型强化学习模仿学习最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。