arXiv:2505.23426cs.LGcs.AI2025-05被引 15

用梯度引导提升扩散策略效率,5步完成高精度控制

Enhanced DACER Algorithm with High Diffusion Efficiency

  • 引入动作梯度场作为辅助目标,每步扩散都有监督信号
  • 仅需5步扩散即在复杂任务上超越现有算法性能
  • 适合实时在线强化学习,尤其擅长多模态控制

由于强大的表达能力,扩散模型在离线强化学习和模仿学习中展现出巨大潜力。扩散演员-评论家带熵调节器(DACER)通过将逆扩散过程作为策略近似器,将这一能力扩展到在线强化学习,实现了顶尖性能。然而,其仍面临核心权衡:更多扩散步数可保证高性能但降低效率,步数减少则性能下降。这成为扩散策略在实时在线强化学习中部署的主要瓶颈。为此,我们提出DACERv2,利用相对于动作的Q梯度场目标作为辅助优化目标,引导每一步扩散过程中的去噪,从而引入中间监督信号,提升单步扩散效率。此外,我们发现Q梯度场与扩散时间步独立的假设与扩散过程特性不符。为此,引入时序加权机制,使模型能在早期有效消除大规模噪声,在后期精细优化输出。在OpenAI Gym基准和多模态任务上的实验表明,相比经典及基于扩散的在线强化学习算法,DACERv2在多数复杂控制环境中仅用5步扩散即实现更高性能,并表现出更强的多模态性。

原文摘要 · Abstract (English)

Due to their expressive capacity, diffusion models have shown great promise in offline RL and imitation learning. Diffusion Actor-Critic with Entropy Regulator (DACER) extended this capability to online RL by using the reverse diffusion process as a policy approximator, achieving state-of-the-art performance. However, it still suffers from a core trade-off: more diffusion steps ensure high performance but reduce efficiency, while fewer steps degrade performance. This remains a major bottleneck for deploying diffusion policies in real-time online RL. To mitigate this, we propose DACERv2, which leverages a Q-gradient field objective with respect to action as an auxiliary optimization target to guide the denoising process at each diffusion step, thereby introducing intermediate supervisory signals that enhance the efficiency of single-step diffusion. Additionally, we observe that the independence of the Q-gradient field from the diffusion time step is inconsistent with the characteristics of the diffusion process. To address this issue, a temporal weighting mechanism is introduced, allowing the model to effectively eliminate large-scale noise during the early stages and refine its outputs in the later stages. Experimental results on OpenAI Gym benchmarks and multimodal tasks demonstrate that, compared with classical and diffusion-based online RL algorithms, DACERv2 achieves higher performance in most complex control environments with only five diffusion steps and shows greater multimodality.

扩散模型强化学习高效策略多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。