为马尔可夫决策过程提供可解释的归因分析,揭示状态与路径的重要性。
Attribution-based Explanations for Markov Decision Processes

- 基于策略合成技术计算状态与路径的归因分数
- 在五个案例中验证了对决策逻辑的可解释性提升
- 适合关注强化学习可解释性的研究者与工程师
归因技术通过为输入分配数值评分来解释AI模型的输出。现有方法主要针对单一时点的静态特征,难以推广到序列决策场景。本文填补这一空白,提出适用于马尔可夫决策过程(MDP)的归因解释方法。我们形式化定义了MDP中归因应表达的内容,聚焦于对单个状态和执行路径的重要性评分。通过利用策略合成技术,即便在存在非确定性的MDP中,也能高效计算这些重要性分数。我们在五个案例研究中评估了该方法,证明其能提供对序列决策代理逻辑的可解释洞察。
原文摘要 · Abstract (English)
Attribution techniques explain the outcome of an AI model by assigning a numerical score to its inputs. So far, these techniques have mainly focused on attributing importance to static input features at a single point in time, and thus fail to generalize to sequential decision-making settings. This paper fills this gap by introducing techniques to generate attribution-based explanations for Markov Decision Processes (MDPs). We give a formal characterization of what attributions should represent in MDPs, focusing on explanations that assign importance scores to both individual states and execution paths. We show how importance scores can be computed by leveraging techniques for strategy synthesis, enabling the efficient computation of these scores despite the non-determinism inherent in an MDP. We evaluate our approach on five case-studies, demonstrating its utility in providing interpretable insights into the logic of sequential decision-making agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。