通过模块化设计提升强化学习模型可解释性,发现导航功能模块的稳定出现。
Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning
- 在策略网络中通过稀疏性和局部性约束诱导功能模块生成。
- 在2D/3D MiniGrid中识别出沿不同轴的独立导航模块。
- 通过权重干预验证模块功能,适合关注RL可解释性的研究者。
可解释性对确保强化学习系统与人类价值观对齐至关重要,但在复杂决策领域仍具挑战。现有方法常试图在神经元或决策节点层面实现可解释性,但难以扩展至大型模型。本文提出在功能模块层面实现可解释性:通过在网络权重中引入稀疏性和局部性约束,促使强化学习策略网络中涌现出功能模块。为检测这些模块,我们开发了一种改进的Louvain算法,采用新型“相关对齐”度量,克服了标准网络分析在神经网络架构中的局限性。在2D和3D MiniGrid环境中应用该方法,揭示了沿不同坐标轴的稳定导航模块,且通过推理前直接干预网络权重,进一步验证了这些功能模块的有效性。
原文摘要 · Abstract (English)
Interpretability is crucial for ensuring RL systems align with human values. However, it remains challenging to achieve in complex decision making domains. Existing methods frequently attempt interpretability at the level of fundamental model units, such as neurons or decision nodes: an approach which scales poorly to large models. Here, we instead propose an approach to interpretability at the level of functional modularity. We show how encouraging sparsity and locality in network weights leads to the emergence of functional modules in RL policy networks. To detect these modules, we develop an extended Louvain algorithm which uses a novel `correlation alignment' metric to overcome the limitations of standard network analysis techniques when applied to neural network architectures. Applying these methods to 2D and 3D MiniGrid environments reveals the consistent emergence of distinct navigational modules for different axes, and we further demonstrate how these functions can be validated through direct interventions on network weights prior to inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。