让智能体在不确定环境下更谨慎决策,降低极端风险。
Online Risk-Averse Planning in POMDPs Using Iterated CVaR Value Function
- 用迭代条件风险价值优化在线规划,动态调节风险偏好。
- 在多个基准测试中,尾部风险显著低于传统方法。
- 适合高风险场景下的机器人、金融等安全敏感应用。
我们研究在部分可观测环境下使用动态风险度量迭代条件风险价值(ICVaR)进行风险敏感规划。提出一种用于ICVaR的策略评估算法,具备不依赖动作空间基数的有限时间性能保证。在此基础上,将三种常用在线规划算法——稀疏采样、带双渐进扩展的粒子滤波树(PFT-DPW)和带观测扩展的部分可观测蒙特卡洛规划(POMCPOW)——扩展为优化ICVaR价值函数而非回报期望。新方法引入风险参数α,当α=1时恢复标准期望规划,α<1则增强风险规避。针对ICVaR稀疏采样,建立了基于风险敏感目标的有限时间性能保证,进而设计了一种专为ICVaR定制的新探索策略。在多个基准POMDP领域上的实验表明,所提出的ICVaR规划器相比风险中立方法显著降低了尾部风险。
原文摘要 · Abstract (English)
We study risk-sensitive planning under partial observability using the dynamic risk measure Iterated Conditional Value-at-Risk (ICVaR). A policy evaluation algorithm for ICVaR is developed with finite-time performance guarantees that do not depend on the cardinality of the action space. Building on this foundation, three widely used online planning algorithms--Sparse Sampling, Particle Filter Trees with Double Progressive Widening (PFT-DPW), and Partially Observable Monte Carlo Planning with Observation Widening (POMCPOW)--are extended to optimize the ICVaR value function rather than the expectation of the return. Our formulations introduce a risk parameter $α$, where $α= 1$ recovers standard expectation-based planning and $α< 1$ induces increasing risk aversion. For ICVaR Sparse Sampling, we establish finite-time performance guarantees under the risk-sensitive objective, which further enable a novel exploration strategy tailored to ICVaR. Experiments on benchmark POMDP domains demonstrate that the proposed ICVaR planners achieve lower tail risk compared to their risk-neutral counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。