arXiv:2509.18147cs.LGcs.AI2025-09

通过追踪概念在卷积层间的演化,揭示CNN的内部推理逻辑。

ConceptFlow: Hierarchical and Fine-grained Concept-Based Explanation for Convolutional Neural Networks

  • 为每个卷积滤波器关联高层语义概念,实现局部语义解释。
  • 构建概念转移矩阵,量化概念在层间传播与转化路径。
  • 适合关注模型可解释性与决策过程分析的研究者。

基于概念的卷积神经网络可解释性旨在对齐模型内部表征与高层语义概念,但现有方法大多忽视单个滤波器的语义角色以及概念在层间动态传播。为此,我们提出ConceptFlow,一种模拟模型内部“思维路径”的概念可解释性框架,通过追踪概念如何在各层中生成与演变来揭示模型推理过程。ConceptFlow包含两个核心组件:(i) 概念注意力,将每个滤波器与相关高层概念关联,实现局部语义解释;(ii) 概念路径,基于概念转移矩阵量化概念在滤波器间的传播与转变。二者共同提供统一、结构化的内部推理视图。实验表明,ConceptFlow能产生语义上合理的模型推理洞察,验证了概念注意力与概念路径在解释决策行为方面的有效性。通过建模分层的概念路径,ConceptFlow深化了对CNN内部逻辑的理解,并支持生成更忠实、更符合人类认知的解释。

原文摘要 · Abstract (English)

Concept-based interpretability for Convolutional Neural Networks (CNNs) aims to align internal model representations with high-level semantic concepts, but existing approaches largely overlook the semantic roles of individual filters and the dynamic propagation of concepts across layers. To address these limitations, we propose ConceptFlow, a concept-based interpretability framework that simulates the internal "thinking path" of a model by tracing how concepts emerge and evolve across layers. ConceptFlow comprises two key components: (i) concept attentions, which associate each filter with relevant high-level concepts to enable localized semantic interpretation, and (ii) conceptual pathways, derived from a concept transition matrix that quantifies how concepts propagate and transform between filters. Together, these components offer a unified and structured view of internal model reasoning. Experimental results demonstrate that ConceptFlow yields semantically meaningful insights into model reasoning, validating the effectiveness of concept attentions and conceptual pathways in explaining decision behavior. By modeling hierarchical conceptual pathways, ConceptFlow provides deeper insight into the internal logic of CNNs and supports the generation of more faithful and human-aligned explanations.

可解释性卷积网络概念追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。