arXiv:2601.22993cs.LGstat.ML2026-01

提出新方法Canary,让强化学习更安全地控制风险。

Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk

  • 用坎特利不等式构建稳定的风险约束估计
  • 在密集成本场景下违规最少且最快达标
  • 适合需要严格安全约束的连续控制任务

我们提出Canary,一种面向风险规避的强化学习方法,用于优化受价值风险(VaR)约束的RL问题。通过应用坎特利不等式,基于收益成本的前两阶矩,获得一个可处理、保守且平滑的VaR约束上界,使约束估计在密集成本环境下仍保持稳定且违反阈值紧致。在此基础上,扩展了约束策略优化(CPO)的信任域框架,为训练过程中的策略改进和约束违反提供最坏情况边界。实验表明,在连续控制安全基准测试中,Canary 最可靠地满足约束,违规次数最少,且最早实现永久性满足,同时在奖励表现上与其他合规基线相当。

原文摘要 · Abstract (English)

We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems. We employ Cantelli's inequality to obtain a tractable, conservative and smooth bound on the VaR constraint based on the first two moments of the cost return. This yields a constraint estimator that remains stable with tight violation thresholds in dense cost regimes. Extending the trust-region framework of the Constrained Policy Optimization (CPO) method, we further provide worst-case bounds for both policy improvement and constraint violation during the training process. Empirically, across continuous-control safety benchmarks, Canary most reliably satisfies its constraint, with the fewest violations and the earliest permanent satisfaction, while remaining reward-competitive with other baselines that also satisfy.

强化学习风险控制安全决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。