arXiv:2409.01245cs.LGcs.AI2024-09被引 1

提出新安全度量EMCC,更好区分连续与偶发危险行为。

Revisiting Safe Exploration in Safe Reinforcement learning

  • 用连续成本步数评估风险严重性,改进传统安全约束
  • 在多种算法中验证,显著提升探索安全性
  • 新增轻量级基准任务,便于快速算法评测

安全强化学习(SafeRL)通过将轨迹的期望成本回报控制在阈值以下来定义安全。然而,该指标无法区分成本累积方式,将罕见严重事件与频繁轻微事件等同处理,可能导致更危险的探索行为。本文提出新度量指标——期望最大连续成本步数(EMCC),通过评估不安全步骤的连续发生情况来衡量风险严重性,特别适用于区分长期与偶发的安全违规。该方法被应用于基于策略和离策略算法中,用于评估其安全探索能力。最后,通过一系列基准测试验证了该度量的有效性,并提出一个轻量级新基准任务,支持快速算法设计与评估。

原文摘要 · Abstract (English)

Safe reinforcement learning (SafeRL) extends standard reinforcement learning with the idea of safety, where safety is typically defined through the constraint of the expected cost return of a trajectory being below a set limit. However, this metric fails to distinguish how costs accrue, treating infrequent severe cost events as equal to frequent mild ones, which can lead to riskier behaviors and result in unsafe exploration. We introduce a new metric, expected maximum consecutive cost steps (EMCC), which addresses safety during training by assessing the severity of unsafe steps based on their consecutive occurrence. This metric is particularly effective for distinguishing between prolonged and occasional safety violations. We apply EMMC in both on- and off-policy algorithm for benchmarking their safe exploration capability. Finally, we validate our metric through a set of benchmarks and propose a new lightweight benchmark task, which allows fast evaluation for algorithm design.

强化学习安全评估风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。