arXiv:2603.22292cs.LGcs.AI2026-03中稿 · the 36th Internati…被引 2

提出预算约束下的安全可达集,让离线强化学习更安全高效

Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning

  • 定义安全条件可达集,解耦奖励与累积成本约束
  • 无需不稳定优化,从固定数据集学出安全策略
  • 在标准测试和海上导航任务中表现优于现有方法

序列决策问题广泛存在于现实应用中,基于马尔可夫决策过程的方法在该领域已取得显著成果。然而,真实任务需权衡奖励最大化与安全约束,二者常冲突,易引发不稳定的极值或对抗性优化。一种有前景的替代方案是安全可达性分析,其预计算一个前向不变的安全状态-动作集,确保智能体从该集合内启动后可无限期保持安全。但多数可达性方法仅处理硬性安全约束,极少研究将可达性扩展至累积成本约束。为此,本文首先定义安全条件可达集,实现奖励最大化与累积安全成本约束的解耦;其次,证明该集合可在不引入不稳定极值或拉格朗日优化的前提下强制执行安全约束,从而提出一种新颖的离线安全强化学习算法,仅从固定数据集学习安全策略而无需环境交互;最后,在标准离线安全强化学习基准及一个真实的海上航行任务上进行实验,结果表明该方法在保持安全性的前提下性能达到或超过当前最优基线。

原文摘要 · Abstract (English)

Sequential decision making using Markov Decision Process underpins many realworld applications. Both model-based and model free methods have achieved strong results in these settings. However, real-world tasks must balance reward maximization with safety constraints, often conflicting objectives, that can lead to unstable min/max, adversarial optimization. A promising alternative is safety reachability analysis, which precomputes a forward-invariant safe state, action set, ensuring that an agent starting inside this set remains safe indefinitely. Yet, most reachability based methods address only hard safety constraints, and little work extends reachability to cumulative cost constraints. To address this, first, we define a safetyconditioned reachability set that decouples reward maximization from cumulative safety cost constraints. Second, we show how this set enforces safety constraints without unstable min/max or Lagrangian optimization, yielding a novel offline safe RL algorithm that learns a safe policy from a fixed dataset without environment interaction. Finally, experiments on standard offline safe RL benchmarks, and a real world maritime navigation task demonstrate that our method matches or outperforms state of the art baselines while maintaining safety.

强化学习安全控制离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。