arXiv:2602.02924cs.LGcs.SY2026-02被引 2

用改进的拉格朗日函数让扩散模型安全地做强化学习。

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?

  • 引入增强拉格朗日项,使扩散模型的训练更稳定。
  • 在多个环境上实现强且稳定的性能表现。
  • 适合需要安全策略的在线强化学习场景。

扩散策略采样使强化学习能表示多模态动作分布,超越次优的单模高斯策略。然而,现有基于扩散的强化学习方法主要关注离线环境中的奖励最大化,对在线环境下的安全性考虑不足。为此,我们提出增强拉格朗日引导扩散(ALGD)算法,一种新型的离线策略安全强化学习方法。通过重新审视优化理论与能量模型,我们发现原始对偶方法的不稳定性源于非凸的拉格朗日景观。在基于扩散的安全强化学习中,拉格朗日可被解释为引导去噪动态的能量函数。出人意料的是,直接使用会同时破坏策略生成与训练过程。ALGD通过引入局部凸化的增强拉格朗日项,实现了策略生成与训练的稳定,且不改变最优策略的分布。理论分析与大量实验表明,ALGD兼具理论基础与实际有效性,在多种环境中均表现出强而稳定的性能。

原文摘要 · Abstract (English)

Diffusion policy sampling enables reinforcement learning (RL) to represent multimodal action distributions beyond suboptimal unimodal Gaussian policies. However, existing diffusion-based RL methods primarily focus on offline settings for reward maximization, with limited consideration of safety in online settings. To address this gap, we propose Augmented Lagrangian-Guided Diffusion (ALGD), a novel algorithm for off-policy safe RL. By revisiting optimization theory and energy-based model, we show that the instability of primal-dual methods arises from the non-convex Lagrangian landscape. In diffusion-based safe RL, the Lagrangian can be interpreted as an energy function guiding the denoising dynamics. Counterintuitively, direct usage destabilizes both policy generation and training. ALGD resolves this issue by introducing an augmented Lagrangian that locally convexifies the energy landscape, yielding a stabilized policy generation and training process without altering the distribution of the optimal policy. Theoretical analysis and extensive experiments demonstrate that ALGD is both theoretically grounded and empirically effective, achieving strong and stable performance across diverse environments.

强化学习扩散模型安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。