在不确定环境下,高效求解智能体能否100%达成目标路径或周期性任务。
Qualitative Analysis of $ω$-Regular Objectives on Robust MDPs
- 通过查询不确定性集设计算法,解决无结构限制的鲁棒MDP中目标可达与奇偶性问题。
- 算法在上千状态规模的典型例子上验证有效,实现概率1的目标保障。
- 适用于需要高可靠性决策的系统,如自动驾驶、工业控制等安全关键场景。
鲁棒马尔可夫决策过程(RMDPs)通过定义可能的转移函数集合,扩展了经典MDP对转移概率不确定性的建模能力。目标是满足某种运行轨迹(无限路径)的集合,其值为智能体在对抗环境中能保证的最大概率。本文研究两类目标:(a) 可达性目标,即最终到达指定状态集;(b) 奇偶性目标,作为ω-正则目标的典型形式。定性分析问题关注目标是否能以概率1确保。本工作在不假设RMDP结构(如单链或非周期性)的前提下,研究可达性与奇偶性目标的定性问题。主要贡献包括:首先,提出基于不确定性集查询的高效算法,解决两类目标的定性问题;其次,在文献中的经典RMDP实例上进行实验,结果表明该方法在数千状态规模下仍具有效性。
原文摘要 · Abstract (English)
Robust Markov Decision Processes (RMDPs) generalize classical MDPs that consider uncertainties in transition probabilities by defining a set of possible transition functions. An objective is a set of runs (or infinite trajectories) of the RMDP, and the value for an objective is the maximal probability that the agent can guarantee against the adversarial environment. We consider (a) reachability objectives, where given a target set of states, the goal is to eventually arrive at one of them; and (b) parity objectives, which are a canonical representation for $ω$-regular objectives. The qualitative analysis problem asks whether the objective can be ensured with probability 1. In this work, we study the qualitative problem for reachability and parity objectives on RMDPs without making any assumption over the structures of the RMDPs, e.g., unichain or aperiodic. Our contributions are twofold. We first present efficient algorithms with oracle access to uncertainty sets that solve qualitative problems of reachability and parity objectives. We then report experimental results demonstrating the effectiveness of our oracle-based approach on classical RMDP examples from the literature scaling up to thousands of states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。