arXiv:2410.05655cs.LG2024-10被引 6

提出安全约束下方差最小的策略评估方法,兼顾高效与安全。

Efficient Policy Evaluation with Safety Constraint for Reinforcement Learning

  • 在安全约束下优化行为策略以最小化评估方差
  • 实验表明该方法同时实现大幅降方差与安全执行
  • 适合需要高可靠性与低风险的强化学习应用

在强化学习中,传统在线策略评估方法常因方差过高且需大量在线数据才能达到所需精度而受限。以往研究尝试通过搜索或设计合适的行为策略来降低评估方差,但忽略了这些策略的安全性——所设计的行为策略缺乏安全保证,可能在在线执行时造成严重损害。本文提出一种在安全约束下最优方差最小化的行为策略。理论上,在确保安全的前提下,该评估方法无偏且方差低于在线策略评估。实验表明,该方法是目前唯一能同时实现显著方差降低与安全约束满足的方法,且在方差减少与执行安全性方面均优于先前方法。

原文摘要 · Abstract (English)

In reinforcement learning, classic on-policy evaluation methods often suffer from high variance and require massive online data to attain the desired accuracy. Previous studies attempt to reduce evaluation variance by searching for or designing proper behavior policies to collect data. However, these approaches ignore the safety of such behavior policies -- the designed behavior policies have no safety guarantee and may lead to severe damage during online executions. In this paper, to address the challenge of reducing variance while ensuring safety simultaneously, we propose an optimal variance-minimizing behavior policy under safety constraints. Theoretically, while ensuring safety constraints, our evaluation method is unbiased and has lower variance than on-policy evaluation. Empirically, our method is the only existing method to achieve both substantial variance reduction and safety constraint satisfaction. Furthermore, we show our method is even superior to previous methods in both variance reduction and execution safety.

强化学习策略评估安全约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。