提出新方法让强化学习更关注极端风险,提升安全决策能力。
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
- 用状态扩展重构风险目标,使每步奖励更密集
- 算法收敛且在离散化下有误差保证
- 适合高安全性要求的强化学习任务
尾部风险度量如静态条件风险价值(CVaR)用于安全关键场景,防止罕见但灾难性事件。与风险中性目标不同,静态CVaR依赖完整轨迹,无法在马尔可夫决策过程中原递归分解。经典方法通过连续变量进行状态扩展,但除非限制于特定值函数类,否则会导致稀疏奖励和退化解。本文提出基于扩展的新静态CVaR形式化,其贝尔曼算子具有:(1) 密集的每步奖励;(2) 在所有有界值函数空间上的压缩性质。基于此理论基础,我们开发了风险规避值迭代与无模型Q-learning算法,依赖离散化扩展状态。进一步提供收敛性保证及离散化引起的近似误差界。实验表明,算法能有效学习对CVaR敏感的策略,实现性能与安全性的良好权衡。
原文摘要 · Abstract (English)
Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral objectives, the static CVaR of the return depends on entire trajectories without admitting a recursive Bellman decomposition in the underlying Markov decision process. A classical resolution relies on state augmentation with a continuous variable. However, unless restricted to a specialized class of admissible value functions, this formulation induces sparse rewards and degenerate fixed points. In this work, we propose a novel formulation of the static CVaR objective based on augmentation. Our alternative approach leads to a Bellman operator with: (1) dense per-step rewards; (2) contracting properties on the full space of bounded value functions. Building on this theoretical foundation, we develop risk-averse value iteration and model-free Q-learning algorithms that rely on discretized augmented states. We further provide convergence guarantees and approximation error bounds due to discretization. Empirical results demonstrate that our algorithms successfully learn CVaR-sensitive policies and achieve effective performance-safety trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。