arXiv:2606.22630cs.LG2026-06

无需真实数据,用伴随匹配高效训练扩散策略

Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching

论文配图:Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching
图 1 · 摘自论文原文
  • 通过伴随匹配实现无仿真训练,避开复杂反向传播
  • 计算开销显著降低,性能媲美传统方法
  • 适合追求高效在线强化学习的工程与研究者

扩散策略最近成为强化学习中表示复杂动作分布的强大范式。然而,在缺乏真实数据的情况下,其在线强化学习应用受限于难以规模化训练的问题,标准优化技术如得分匹配无法直接适用。本文提出一种高效算法,利用随机最优控制最新进展,基于伴随匹配优化扩散策略,实现无需仿真的训练,避免显式似然估计和昂贵的扩散过程反向传播。此外,我们提出若干扩展,提升方法在实际场景中的鲁棒性与稳定性。实验表明,该方法在保持竞争力表现的同时显著降低计算开销,使扩散策略更适用于在线强化学习场景。

原文摘要 · Abstract (English)

Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online RL remains limited by the challenge of scalable training in the absence of ground-truth data, where standard optimization techniques such as score matching are not directly applicable. In this work, we introduce a highly efficient algorithm for optimizing diffusion policies by leveraging recent advances in stochastic optimal control. Our approach is based on adjoint matching, which enables simulation-free training and circumvents the need for explicit likelihood estimation or costly backpropagation through the diffusion process. Furthermore, we propose several extensions that improve the robustness and stability of the method in practical settings. Empirical results demonstrate that our approach achieves competitive performance while significantly reducing computational overhead, making diffusion policies more viable for online RL scenarios.

扩散模型强化学习高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。