arXiv:2502.13731cs.AI2025-02被引 3

提出新方法在马尔可夫决策过程中计算鲁棒反事实概率边界。

Robust Counterfactual Inference in Markov Decision Processes

  • 基于所有相容因果模型,给出反事实转移概率的紧界。
  • 无需求解大规模优化,提供闭式表达式,计算高效。
  • 可生成最坏情况下的鲁棒策略,适合不确定性场景。

本文针对现有马尔可夫决策过程(MDP)反事实推断方法的关键局限:当前方法依赖特定因果模型以使反事实可识别,但通常存在多个与观测和干预分布一致的因果模型,导致不同反事实分布,固定单一模型会削弱推断的有效性与实用性。为此,我们提出一种新的非参数方法,计算所有相容因果模型下反事实转移概率的紧界。不同于以往需求解变量随MDP规模呈指数增长的复杂优化问题,本方法提供闭式表达式,实现高效率、可扩展的计算。构建区间反事实MDP后,进一步识别在不确定概率下最坏情况奖励最优的鲁棒反事实策略。在多个案例研究中验证,该方法显著优于现有方法的鲁棒性。

原文摘要 · Abstract (English)

This paper addresses a key limitation in existing counterfactual inference methods for Markov Decision Processes (MDPs). Current approaches assume a specific causal model to make counterfactuals identifiable. However, there are usually many causal models that align with the observational and interventional distributions of an MDP, each yielding different counterfactual distributions, so fixing a particular causal model limits the validity (and usefulness) of counterfactual inference. We propose a novel non-parametric approach that computes tight bounds on counterfactual transition probabilities across all compatible causal models. Unlike previous methods that require solving prohibitively large optimisation problems (with variables that grow exponentially in the size of the MDP), our approach provides closed-form expressions for these bounds, making computation highly efficient and scalable for non-trivial MDPs. Once such an interval counterfactual MDP is constructed, our method identifies robust counterfactual policies that optimise the worst-case reward w.r.t. the uncertain interval MDP probabilities. We evaluate our method on various case studies, demonstrating improved robustness over existing methods.

反事实推理强化学习鲁棒性因果建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。