arXiv:2606.29681cs.AIcs.SY2026-06中稿 · UAI2026 as oral pr…

提出新方法,高效识别马尔可夫决策过程中的概率因果状态。

Sample-Efficient Learning of Probabilistic Causes for Reachability in Markov Decision Processes with Probabilistic Guarantees

论文配图:Sample-Efficient Learning of Probabilistic Causes for Reachability in Markov Decision Processes with Probabilistic Guarantees
图 1 · 摘自论文原文
  • 用重启机制将因果检测转为两个条件可达性查询
  • 无需原始可达概率即可完成因果分析,支持学习
  • 算法可渐进分类状态,适合在线推理场景

马尔可夫决策过程(MDP)的概率模型检测能提供定量保证,但难以解释不良结果的原因。概率提升(PR)因果分析通过识别那些访问会提高目标状态可达概率的状态来解决此问题。现有方法依赖不适用于学习的MDP修改:从转移样本中检测条件与无条件可达概率的差距困难,且需已知原MDP的可达概率,当转移概率未知时无法使用。本文研究未知MDP,提出一种具有概率保证的PR因果识别学习方法。核心是基于重启的MDP修改,将因果检测转化为两个无需原MDP可达值的条件可达性查询。我们证明了方法正确性,建立了样本复杂度界,并设计了一种基于双向值迭代的任意时学习-验证算法,逐步将状态分类为因果、非因果或待定。在两个基准测试上的实验表明,该方法能可靠且快速地识别出PR因果状态。

原文摘要 · Abstract (English)

Probabilistic model checking for Markov decision processes (MDPs) provides quantitative guarantees, but often offers limited insight into why undesired outcomes occur. Probability-raising (PR) causality addresses this by identifying states whose visitation increases the probability of reaching designated states. Existing PR-cause identification methods, however, use MDP modifications not well-suited for learning: the gap between conditional and unconditional reachability probabilities can be hard to detect from transition samples, and construction requires reachability probabilities of the MDP, which are unavailable when transition probabilities are unknown. We study unknown MDPs and propose a learning approach with probabilistic guarantees for PR-cause identification. Our key ingredient is a restart-based MDP modification that reduces PR-cause checking to two conditional reachability queries without using reachability values of the original MDP. We prove correctness, establish sample-complexity bounds, and develop an anytime learning-and-checking algorithm based on two-sided value iteration that progressively classifies states as causal, non-causal, or undecided. Experiments on two benchmarks demonstrate reliable and fast identification of PR causes.

因果推断MDP强化学习概率保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。