arXiv:2605.02777cs.LGcs.AI2026-05

提出新方法让强化学习安全适应动态成本限制,提升安全性和奖励表现。

Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning

论文配图:Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning
图 1 · 摘自论文原文
  • 将安全轨迹生成视为受限区域采样,分离成本与奖励引导。
  • 在DSRL基准上94.7%任务满足约束,21个任务中奖励最高。
  • 适合需要动态安全控制的机器人、自动驾驶等部署场景。

离线安全强化学习常需策略在部署时适应随回合变化或单回合内变动的安全预算。尽管基于扩散的规划器支持灵活轨迹生成,现有引导机制常将奖励提升与约束满足视为竞争目标,导致成本限制下安全性不可靠。本文将自适应安全轨迹生成重新理解为从受限轨迹分布中采样:预算限定轨迹范围,奖励在该范围内塑造偏好。据此提出安全解耦引导扩散(SDGD),通过条件分类器自由引导以成本上限为依据偏向采样满足限制的轨迹,同时用奖励梯度引导优化回报。由于直接奖励引导可能在提升回报的同时引导样本进入高累积成本路径,我们引入可行轨迹重标注(FTR)重构奖励目标,抑制此类方向。进一步提供一阶采样时间分析,证明在前缀可恢复对齐条件下FTR能抑制奖励引发的成本漂移。在DSRL基准上的大量实验表明,SDGD在基线中安全性最优,在38个任务中有36个满足约束(94.7%),且在21个任务中获得安全方法中的最高回报。

原文摘要 · Abstract (English)

Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a single episode. While diffusion-based planners enable flexible trajectory generation, existing guidance schemes often treat reward improvement and constraint satisfaction as competing gradient objectives, which can lead to unreliable safety compliance under cost limits. We reinterpret adaptive safe trajectory generation as sampling from a constrained trajectory distribution, where the budget restricts the trajectory region, and reward shapes preferences within that region. This perspective motivates Safe Decoupled Guidance Diffusion (SDGD), which conditions classifier-free guidance on the cost limit to bias sampling toward trajectories satisfying the specified limit, while using reward-gradient guidance to refine trajectories for higher return. Because direct reward guidance can increase return while also steering samples toward trajectories with higher cumulative cost, we introduce Feasible Trajectory Relabeling (FTR) to reshape reward targets and discourage such directions. We further provide a first-order sampling-time analysis showing that FTR suppresses reward-induced cost drift under a prefix-restorative alignment condition. Extensive evaluations on the DSRL benchmark show that SDGD achieves the strongest safety compliance among baselines, satisfying the constraint on 94.7% of tasks (36/38), while obtaining the highest reward among safe methods on 21 tasks.

强化学习安全控制扩散模型离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。