arXiv:2603.23889cs.LGcs.RO2026-03中稿 · ICLR被引 2

提出新算法让强化学习在保证安全前提下高效探索。

Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration

  • 用约束乐观探索策略解决奖励与成本的梯度冲突。
  • 通过截断分位数批评家稳定成本价值学习并控制风险。
  • 适合自动驾驶等高安全要求场景,兼顾效率与安全性。

当安全被定义为累积成本的上限时,安全强化学习旨在学习最大化回报且满足成本约束的策略。离线安全强化学习方法虽具高样本效率,但因探索无成本意识及累积成本估计偏差,常导致约束违反。为此,本文提出约束乐观探索Q学习(COX-Q),结合成本受限在线探索与保守离线分布值学习。首先,提出一种新型成本约束乐观探索策略,解决动作空间中奖励与成本间的梯度冲突,并自适应调整信任区域以控制训练成本。其次,采用截断分位数批评家稳定成本值学习,同时量化认知不确定性以引导探索。在安全速度控制、安全导航和自动驾驶任务上的实验表明,COX-Q实现了高样本效率、具有竞争力的测试安全性能,并有效控制数据采集成本。结果表明,COX-Q是安全关键应用中一种有前景的强化学习方法。

原文摘要 · Abstract (English)

When safety is formulated as a limit of cumulative cost, safe reinforcement learning (RL) aims to learn policies that maximize return subject to the cost constraint in data collection and deployment. Off-policy safe RL methods, although offering high sample efficiency, suffer from constraint violations due to cost-agnostic exploration and estimation bias in cumulative cost. To address this issue, we propose Constrained Optimistic eXploration Q-learning (COX-Q), an off-policy safe RL algorithm that integrates cost-bounded online exploration and conservative offline distributional value learning. First, we introduce a novel cost-constrained optimistic exploration strategy that resolves gradient conflicts between reward and cost in the action space and adaptively adjusts the trust region to control the training cost. Second, we adopt truncated quantile critics to stabilize the cost value learning. Quantile critics also quantify epistemic uncertainty to guide exploration. Experiments on safe velocity, safe navigation, and autonomous driving tasks demonstrate that COX-Q achieves high sample efficiency, competitive test safety performance, and controlled data collection cost. The results highlight COX-Q as a promising RL method for safety-critical applications.

强化学习安全控制离线学习自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。