让深度强化学习模型的神经元行为可解释,自动关联决策逻辑。
Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning
- 用逻辑公式自动关联神经元激活与语义概念
- 通过值敏感离散化提取关键决策边界概念
- 适合需要信任AI决策的高风险场景研究者
深度强化学习在复杂控制任务中表现优异,但其策略或价值网络仍缺乏可解释性,影响在高风险场景中的可信度。现有基于概念的方法在计算机视觉中有效,但在连续状态空间的DRL中受限于缺乏预定义语义概念。本文提出一种新型概念解释框架,实现神经元级别的细粒度可解释性。不同于依赖人工特征工程的方法,该框架自动将神经元激活对齐到由语义谓词构成的逻辑公式。为连接连续信号与符号推理,引入值敏感离散化机制,将原始状态特征转换为可解释的原子概念,确保解释词汇捕捉与智能体价值评估相关的战略决策边界。通过组合这些可解释概念并匹配神经元行为,获得网络内部表征的明确解释。在连续与离散环境上的实验表明,该方法能有效识别有意义的决策模式,提供符合人类直觉的忠实解释。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) has successfully addressed many complex control problems. However, the neural networks representing policies or values remain opaque, undermining trust in high-stakes applications. While concept-based methods have shown promise in deciphering internal representations in computer vision, applying them to DRL is impeded by the absence of pre-defined semantic concepts in continuous state spaces. In this work, we propose a novel concept-based explanation framework designed to provide fine-grained, neuron-level insights into DRL models. Unlike previous approaches that rely on manual feature engineering, our framework automatically aligns neuron activations with logical formulas composed of semantic predicates. To bridge the gap between continuous signals and symbolic reasoning, we introduce a value-sensitive discretization mechanism that transforms raw state features into interpretable atomic concepts. This ensures that the vocabulary used for explanation captures strategic decision boundaries relevant to the agent's value assessment. By composing these interpretable concepts and matching them with neuron behaviors, we derive explicit explanations for the network's internal representations. Experimental results on both continuous and discrete environments demonstrate that our method effectively identifies meaningful decision-making patterns, offering faithful explanations that align with human intuition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。