让强化学习更安全:用风险感知约束提升系统鲁棒性
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
- 用优化确定等价物建模每阶段风险,兼顾收益与时间鲁棒性
- 理论保证原约束问题等价,算法可嵌入PPO等主流框架
- 适合高风险场景如自动驾驶、医疗决策,关注极端风险的读者必看
约束优化是处理强化学习中冲突目标的通用框架。多数方法以期望累积回报表达目标和约束,但忽略奖励分布尾部的风险事件,难以满足高风险应用需求。本文提出一种风险感知的约束强化学习框架,通过优化确定等价物(OCEs)实现收益值与时间上的联合阶段鲁棒性。在适当的约束资格条件下,该框架在参数化强拉格朗日对偶框架下保证与原问题精确等价,并提供可嵌入标准RL求解器(如PPO)的简单算法流程。最后,在常见假设下证明了算法收敛性,并通过多个数值实验验证了方法的风险感知特性。
原文摘要 · Abstract (English)
Constrained optimization provides a common framework for dealing with conflicting objectives in reinforcement learning (RL). In most of these settings, the objectives (and constraints) are expressed though the expected accumulated reward. However, this formulation neglects risky or even possibly catastrophic events at the tails of the reward distribution, and is often insufficient for high-stakes applications in which the risk involved in outliers is critical. In this work, we propose a framework for risk-aware constrained RL, which exhibits per-stage robustness properties jointly in reward values and time using optimized certainty equivalents (OCEs). Our framework ensures an exact equivalent to the original constrained problem within a parameterized strong Lagrangian duality framework under appropriate constraint qualifications, and yields a simple algorithmic recipe which can be wrapped around standard RL solvers, such as PPO. Lastly, we establish the convergence of the proposed algorithm under common assumptions, and verify the risk-aware properties of our approach through several numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。