揭示了马尔可夫决策中静态CVaR评估的固有局限性
On the Fundamental Limitations of Dual Static CVaR Decompositions in Markov Decision Processes
- 通过风险分配一致性约束分析评估误差根源
- 发现双分解下最优策略无法统一适用于所有风险水平
- 证明了双重静态CVaR分解存在根本性限制
最近研究表明,基于对偶形式的动态规划方法在求解马尔可夫决策过程(MDPs)中静态CVaR最优策略时可能失效,但其根本原因尚不明确。本文转向更简单的策略评估任务,将策略的静态CVaR评估转化为两个不同的最小化问题。我们引入一组“风险分配一致性约束”,只有当这些约束有交集时,两者的解才一致。我们证明,当这些约束无交集时,会导致先前观察到的评估误差。将评估误差量化为“CVaR评估差距”,并表明基于对偶的CVaR DP优化所得到的策略具有非零的评估差距。最后,借助风险分配约束视角,证明在对偶CVaR分解中寻找单一普遍最优策略是根本受限的,并识别出一个在所有初始风险水平下均无统一最优策略的MDP实例。
原文摘要 · Abstract (English)
It was recently shown that dynamic programming (DP) methods for finding static CVaR-optimal policies in Markov Decision Processes (MDPs) can fail when based on the dual formulation, yet the root cause of this failure remains unclear. We expand on these findings by shifting focus from policy optimization to the seemingly simpler task of policy evaluation. We show that evaluating the static CVaR of a given policy can be framed as two distinct minimization problems. We introduce a set of ``risk-assignment consistency constraints'' that must be satisfied for their solutions to match and we demonstrate that an empty intersection of these constraints is the source of previously observed evaluation errors. Quantifying the evaluation error as the \emph{CVaR evaluation gap}, we demonstrate that the issues observed when optimizing over the dual-based CVaR DP are explained by the returned policy having a non-zero CVaR evaluation gap. Finally, we leverage our proposed risk-assignment constraints perspective to prove that the search for a single, uniformly optimal policy on the dual CVaR decomposition is fundamentally limited, identifying an MDP where no single policy can be optimal across all initial risk levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。