用逻辑编程解析强化学习决策,量化解释性并揭示行为演化规律。
Explaining Reinforcement Learning Agents via Inductive Logic Programming

- 通过归纳逻辑编程提取策略的符号规则,实现可读性表达。
- 提出激活率、语义距离等指标,精准衡量规则与行为的一致性。
- 适用于多智能体协作分析,助力策略迁移与泛化研究。
可解释强化学习(XRL)旨在提升强化学习策略的透明度与可理解性,这对安全关键及以人为中心的应用至关重要。然而,现有方法多依赖用户研究,针对特定受众且缺乏统一评估标准。逻辑驱动的XAI方法能提供紧凑、人类可读的决策抽象,但其解释程度的系统性量化仍属开放问题。本文提出面向规划的客观评估指标,用于强化学习中的策略可解释性。结合归纳逻辑编程(ILP)提取策略的符号表示,定义激活率、特征覆盖率、句法距离与语义距离等新指标,量化规则与智能体行为的对齐程度、特征在决策中的作用,以及训练过程与跨智能体间策略的演化。在多种强化学习领域实验表明,该方法不仅揭示动作特异性学习动态,超越全局回报;还能提供细粒度的领域特征洞察,突破传统全局重要性估计;更可识别多智能体中的协作、专业化与适应模式,并为动作特异性策略的迁移与泛化提供关键洞见。
原文摘要 · Abstract (English)
Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific audience and lacking shared evaluation metrics. On the other hand, logic-based approaches within eXplainable Artificial Intelligence (XAI) provide compact, human-readable abstractions of decision-making. However, the systematic quantification of the explainability degree of logical representations remains an open problem. This work aims to advance the state of the art in XRL by introducing objective and planning-oriented metrics for policy explainability in RL settings. At the same time, it contributes to the field of logic for XAI by providing a principled way to quantify the explainability of logical rules, moving beyond common-sense assessments and simple propositional fragments. We employ Inductive Logic Programming (ILP) to extract symbolic representations of RL policies and define a novel set of explainability metrics, including activation rate, feature coverage, syntactic distance and semantic distance. These metrics quantify alignment between symbolic rules and agent behavior, the role of features in decision-making, and the evolution of policies during training and across agents in single and multi-agent RL. Experiments across different RL domains show that the proposed metrics highlight action-specific learning dynamics beyond global return, provide fine-grained insights into domain features beyond classical approaches for global feature importance estimation, and uncover coordination, specialization, and adaptation patterns in MARL. Moreover, they provide crucial insights for the transfer and generalization of action-specific policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。