arXiv:2508.07790cs.AIcs.LO2025-08AAAI

为鲁棒MDP设计更优策略选择,兼顾最坏情况与一般情况表现。

Best-Effort Policies for Robust Markov Decision Processes

  • 在最坏转移概率下优化的同时,最大化一般情况下的期望回报
  • 证明最优鲁棒最佳努力策略存在且可计算,额外开销可控
  • 适合需在不确定环境中平衡安全与性能的决策系统

我们研究了具有转移概率集合的马尔可夫决策过程(即鲁棒MDP,RMDP)的通用情形。标准目标是寻找在对抗性转移概率选择下期望回报最大的策略。当不确定性在状态间独立(称为s-矩形性)时,可通过鲁棒值迭代高效求解最优鲁棒策略。然而,可能存在多个最优鲁棒策略,它们在最坏情况下等价,但在非对抗性转移概率下具有不同的期望回报。为此,我们提出一种改进的策略选择准则,受博弈论中支配与最佳努力概念启发。不只追求最坏情况下的最大回报,还要求在非完全对抗的转移概率下实现最大期望回报。这类策略称为最优鲁棒最佳努力(ORBE)策略。我们证明了ORBE策略始终存在,刻画其结构,并给出一个计算算法,其额外开销相对于标准鲁棒值迭代可接受。该方法为最优鲁棒策略提供了合理的抉择依据。数值实验验证了方法的可行性。

原文摘要 · Abstract (English)

We study the common generalization of Markov decision processes (MDPs) with sets of transition probabilities, known as robust MDPs (RMDPs). A standard goal in RMDPs is to compute a policy that maximizes the expected return under an adversarial choice of the transition probabilities. If the uncertainty in the probabilities is independent between the states, known as s-rectangularity, such optimal robust policies can be computed efficiently using robust value iteration. However, there might still be multiple optimal robust policies, which, while equivalent with respect to the worst-case, reflect different expected returns under non-adversarial choices of the transition probabilities. Hence, we propose a refined policy selection criterion for RMDPs, drawing inspiration from the notions of dominance and best-effort in game theory. Instead of seeking a policy that only maximizes the worst-case expected return, we additionally require the policy to achieve a maximal expected return under different (i.e., not fully adversarial) transition probabilities. We call such a policy an optimal robust best-effort (ORBE) policy. We prove that ORBE policies always exist, characterize their structure, and present an algorithm to compute them with a manageable overhead compared to standard robust value iteration. ORBE policies offer a principled tie-breaker among optimal robust policies. Numerical experiments show the feasibility of our approach.

强化学习鲁棒控制决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。