arXiv:2605.01339cs.LG2026-05

用参数化模型捕捉马尔可夫决策过程的依赖性不确定性,提升鲁棒策略学习精度

Robust Parameter Learning for Uncertain MDPs

  • 将转移概率建模为参数表达式,显式捕获状态转移间的依赖关系
  • 通过统计置信度投影到参数空间,获得更紧致的不确定性估计
  • 适合需要高鲁棒性验证的强化学习系统设计与安全关键场景

基于学习的未知马尔可夫决策过程(MDP)验证方法常采用不确定MDP模型,通过置信区间刻画转移不确定性,并合成对不确定性具有鲁棒性的策略。然而,传统方法通常独立量化每个转移概率的不确定性,忽略了因共享潜在变量导致的依赖关系。本文提出使用参数化MDP(pMDP)来学习此类模型,其中转移概率是参数的函数。我们将经验转移频率的统计不确定性投影到pMDP的参数空间中,生成一个满足大概率正确(PAC)的不确定性模型,该模型保留了转移之间的代数依赖性。由于这类模型求解算法上具有挑战性,我们提出了层次化的、保真的多面体外逼近方法来近似诱导出的置信集。我们在实验中实现了该方法并进行了评估,结果表明其不确定性估计显著优于传统的区间基不确定MDP学习技术。

原文摘要 · Abstract (English)

Learning-based approaches to verifying unknown Markov decision processes (MDPs) often employ uncertain MDPs. These models use, for example, confidence intervals to capture transition uncertainty and allow synthesis of policies that are robust to this uncertainty. However, this approach typically quantifies uncertainty independently for individual transition probabilities, ignoring dependencies due to shared latent quantities. We propose to learn such models using parametric MDPs (pMDPs), where transition probabilities are expressions over a set of parameters. We project statistical uncertainty from empirical transition frequencies onto the pMDP's parameter space, yielding a probably approximately correct (PAC) uncertainty model for the underlying MDP that respects the algebraic dependencies between transitions. The resulting models are algorithmically challenging to solve, so we propose a hierarchy of sound polytopic outer approximations of the induced confidence set. We implement and evaluate our approach, demonstrating substantially tighter uncertainty estimates than classical interval-based uncertain MDP learning techniques.

强化学习不确定建模鲁棒决策参数化MDP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。