研究约束马尔可夫决策过程初始分布变化对解的影响。
Initial Distribution Sensitivity of Constrained Markov Decision Processes
- 通过对偶分析与线性规划扰动,推导最优值随初始分布的变化边界。
- 给出策略因初始分布未知带来的损失上限(后悔界)。
- 适合关注鲁棒决策与初始条件敏感性的研究者。
约束马尔可夫决策过程(CMDPs)由于在不同初始状态分布下不存在统一最优策略,求解复杂度远高于标准MDP,一旦初始分布改变就必须重新求解。本文分析了最优值随初始分布变化的敏感性,利用CMDP的对偶分析和线性规划扰动理论,推导出该变化的上界。进一步说明这些边界可用于评估给定策略因初始分布未知而产生的后悔(regret)。研究揭示了初始分布不确定性对策略性能的影响机制。
原文摘要 · Abstract (English)
Constrained Markov Decision Processes (CMDPs) are notably more complex to solve than standard MDPs due to the absence of universally optimal policies across all initial state distributions. This necessitates re-solving the CMDP whenever the initial distribution changes. In this work, we analyze how the optimal value of CMDPs varies with different initial distributions, deriving bounds on these variations using duality analysis of CMDPs and perturbation analysis in linear programming. Moreover, we show how such bounds can be used to analyze the regret of a given policy due to unknown variations of the initial distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。