用扩散模型提升安全约束下的离线强化学习性能与效率
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
- 用扩散模型提取行为策略,简化后实现快速推理
- 通过梯度调整确保安全约束满足率超90%,奖励表现更优
- 适合需高安全性的机器人控制等真实场景应用
约束强化学习旨在安全约束下获得高性能策略。本文聚焦离线设置,即智能体仅依赖固定数据集——这在真实任务中可避免危险探索。提出扩散正则化约束离线强化学习(DRCORL):首先利用扩散模型从离线数据中捕捉行为策略,再提取简化策略以实现高效推理;进一步采用梯度操控实现安全适应,平衡奖励目标与约束满足。该方法充分利用高质量离线数据并融入安全要求。实验表明,DRCORL 在机器人学习任务中实现可靠的安全性、快速推理和强奖励表现。相比现有安全离线强化学习方法,其始终满足成本限制,且在相同超参数下表现更优,具备实际应用潜力。
原文摘要 · Abstract (English)
Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent has only a fixed dataset -- common in realistic tasks to prevent unsafe exploration. To address this, we propose Diffusion-Regularized Constrained Offline Reinforcement Learning (DRCORL), which first uses a diffusion model to capture the behavioral policy from offline data and then extracts a simplified policy to enable efficient inference. We further apply gradient manipulation for safety adaptation, balancing the reward objective and constraint satisfaction. This approach leverages high-quality offline data while incorporating safety requirements. Empirical results show that DRCORL achieves reliable safety performance, fast inference, and strong reward outcomes across robot learning tasks. Compared to existing safe offline RL methods, it consistently meets cost limits and performs well with the same hyperparameters, indicating practical applicability in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。