基于因果关系的抽象方法,让复杂决策模型更易求解。
Property-driven Causal Abstractions for Markov Decision Processes

- 用状态变量间的因果关系识别可合并的状态
- 在多个基准测试中生成小规模抽象,逼近最优策略
- 适合处理大规模决策问题的研究者与工程师
马尔可夫决策过程(MDP)广泛用于建模决策问题,通常在因子化状态空间中通过状态变量及其取值定义。状态数量的指数级增长使许多推理任务变得困难。抽象技术能缩减MDP规模,缓解可扩展性问题。本文提出一种针对因子化MDP的因果性概念及新型属性驱动的因果抽象方法,保留原模型的多种特性。该方法基于状态变量谓词间的因果关系,识别出因相同原因满足或违反特定抽象属性的状态。我们从理论和实证上比较了不同模型类型(如MDP、区间MDP、随机博弈)下的多种因果抽象方法。实验表明:在多个标准基准上,该方法生成的小型抽象可计算出接近原MDP最优的策略;且其因果抽象常能推广至相关的大规模MDP模型。
原文摘要 · Abstract (English)
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs challenging. Abstractions are promising techniques to reduce MDPs and thus mitigate scalability issues. In this work, we introduce a notion of causality on factored MDPs and a novel property-driven causal abstraction technique that retains many characteristics of the original MDP model. For this, we rely on causal relations over state variable predicates and identify those states that share the same reasons for fulfilling or violating a given abstraction property. We theoretically and empirically compare various causal MDP abstractions using different model types such as MDPs, interval MDPs, or stochastic games. Our evaluation demonstrates the potential of our approach: For several standard benchmarks, we obtain small abstractions that allow us to compute near-optimal policies for the original MDP. Furthermore, our causal abstractions often generalize to related large-scale MDP models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。