通过神经元分组与功能交互,提升模型可解释性。
NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional Interactions
- 从单个神经元转向神经元分组分析,捕捉功能关联
- 自动识别核心神经元并聚类,构建分层交互电路
- 适用于图像调试与概念自动标注,降低计算开销
理解神经网络内部机制对提升模型性能和可解释性至关重要。当前研究多聚焦于单个神经元与模型最终预测之间的关联,但当神经元编码多个无关特征时,难以解释模型内部运作。本文提出新框架 NeurFlow,将关注点从单个神经元转向神经元分组,从神经元-输出关系转向神经元间功能交互。该自动化框架首先基于共享功能关系识别核心神经元并聚类成组,实现对网络内部过程更连贯、可解释的刻画。该方法支持构建跨层神经元交互的分层电路,提升可解释性的同时降低计算成本。大量实证研究验证了 NeurFlow 的保真度。此外,我们展示了其在图像调试与自动概念标注等实际应用中的有效性,凸显其推动神经网络可解释性发展的潜力。
原文摘要 · Abstract (English)
Understanding the inner workings of neural networks is essential for enhancing model performance and interpretability. Current research predominantly focuses on examining the connection between individual neurons and the model's final predictions. Which suffers from challenges in interpreting the internal workings of the model, particularly when neurons encode multiple unrelated features. In this paper, we propose a novel framework that transitions the focus from analyzing individual neurons to investigating groups of neurons, shifting the emphasis from neuron-output relationships to functional interaction between neurons. Our automated framework, NeurFlow, first identifies core neurons and clusters them into groups based on shared functional relationships, enabling a more coherent and interpretable view of the network's internal processes. This approach facilitates the construction of a hierarchical circuit representing neuron interactions across layers, thus improving interpretability while reducing computational costs. Our extensive empirical studies validate the fidelity of our proposed NeurFlow. Additionally, we showcase its utility in practical applications such as image debugging and automatic concept labeling, thereby highlighting its potential to advance the field of neural network explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。