arXiv:2601.12008cs.LG2026-01ICML被引 3

用极端值理论提升强化学习安全性,减少罕见高风险事件。

Extreme Value Policy Optimization for Safe Reinforcement Learning

  • 基于极值理论优化极端成本样本,捕捉尾部风险。
  • 实验显示约束违反概率显著降低,且优于传统期望方法。
  • 适合高安全要求场景,如自动驾驶、医疗决策。

在真实世界应用中,确保强化学习(RL)的安全性至关重要。约束强化学习(CRL)通过在预定义约束下最大化回报来应对这一挑战,通常将约束形式化为预期累积成本。然而,基于期望的约束忽略了尾部分布中罕见但高影响的极端事件(如黑天鹅事件),可能导致严重违规。为此,我们提出极端值策略优化(EVO)算法,利用极值理论(EVT)建模并利用极端奖励与成本样本,减少约束违反。EVO引入极端分位数优化目标,显式捕获成本尾部分布中的极端样本;同时提出一种极端优先重放机制,放大罕见但高影响极端样本的学习信号。理论上,我们建立了策略更新期间预期约束违反的上界,保证在零违规分位数水平上的严格约束满足。进一步表明,EVO相较于期望型方法具有更低的约束违反概率,且方差低于分位数回归方法。大量实验显示,EVO在训练过程中显著减少约束违反,同时保持与基线相当的策略性能。

原文摘要 · Abstract (English)

Ensuring safety is a critical challenge in applying Reinforcement Learning (RL) to real-world scenarios. Constrained Reinforcement Learning (CRL) addresses this by maximizing returns under predefined constraints, typically formulated as the expected cumulative cost. However, expectation-based constraints overlook rare but high-impact extreme value events in the tail distribution, such as black swan incidents, which can lead to severe constraint violations. To address this issue, we propose the Extreme Value policy Optimization (EVO) algorithm, leveraging Extreme Value Theory (EVT) to model and exploit extreme reward and cost samples, reducing constraint violations. EVO introduces an extreme quantile optimization objective to explicitly capture extreme samples in the cost tail distribution. Additionally, we propose an extreme prioritization mechanism during replay, amplifying the learning signal from rare but high-impact extreme samples. Theoretically, we establish upper bounds on expected constraint violations during policy updates, guaranteeing strict constraint satisfaction at a zero-violation quantile level. Further, we demonstrate that EVO achieves a lower probability of constraint violations than expectation-based methods and exhibits lower variance than quantile regression methods. Extensive experiments show that EVO significantly reduces constraint violations during training while maintaining competitive policy performance compared to baselines.

强化学习安全约束极值理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。