arXiv:2607.03903cs.LG2026-07

用扩散模型生成安全策略,实现多任务离线强化学习的高效与可靠。

CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

论文配图:CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning
图 1 · 摘自论文原文
  • 将多任务安全强化学习转化为条件生成问题,利用扩散模型建模
  • 支持不同成本约束且无需重训练,避免分布外动作的错误
  • 适合高风险场景下的多任务智能决策,如机器人控制、自动驾驶

多任务离线安全强化学习旨在从多个任务的离线数据中学习共享的最优安全策略,为高风险、交互成本高的多任务场景提供有效解决方案。然而,多任务、安全约束和分布外(OOD)动作带来的三重挑战使现有方法难以在保障安全的同时最大化回报。本文提出条件扩散模型与上下文提示(CDCP),首先重新审视多任务决策与控制中的需求与挑战,确立多任务离线安全RL目标。随后,将多任务约束优化问题转化为条件生成问题,采用扩散模型建模。在此基础上,设计无分类器引导的成本约束策略,灵活施加成本限制,并通过监督学习消除分布外动作的外推误差。同时引入新颖的上下文提示方法,提升多任务表征精度与对未见任务的适应性。还提出梯度损失同步策略,缓解梯度干扰,提升训练稳定性。大量实验表明,CDCP在多任务场景下性能与安全性均优于当前最先进基线方法,可在不重新训练的情况下满足不同成本约束,为多任务安全强化学习提供更灵活的解决方案。

原文摘要 · Abstract (English)

Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks. This paradigm provides an effective means for the widespread application of RL in multi-task scenarios with high risk and interaction costs. However, the triple challenges of multi-tasking, safety constraints, and out-of-distribution (OOD) actions pose a significant hurdle for existing methods to ensure safety while maximizing reward returns. In this work, we propose a Conditional Diffusion model with Contextual Prompts (CDCP) to address these challenges. Concretely, we first rethink the requirements and challenges in current multi-task decision-making and control scenarios and establish the objectives of multi-task offline safe RL. Subsequently, we transform the multi-task constrained optimization problem into a conditional generation problem using the diffusion model. Based on this, we design a classifier-free guided cost-constraint strategy to provide flexible cost constraints and eliminate extrapolation errors from OOD actions via supervised learning. Additionally, we introduce a novel contextual prompting method to enhance multi-task representation accuracy and adaptability to unseen tasks. A gradient loss synchronization strategy is also introduced to eliminate gradient interference, improving training stability. Finally, extensive experiments demonstrate that the CDCP algorithm exhibits higher performance and safety in multi-task scenarios than the current state-of-the-art baseline methods. It meets different cost constraints without further training, providing a more flexible cost-constraint solution for the multi-task safe RL.

强化学习扩散模型安全决策多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。