用知识图谱解析模型如何逐步形成和传播语义概念。
Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs
- 构建结构化知识图谱,追踪高阶语义概念在模型各层的演变与交互。
- 揭示隐藏的神经通路和信息流动,识别模型决策中的虚假关联。
- 工具化展示,适合研究模型偏差与泛化能力的开发者与研究人员。
传统概念解释方法多聚焦于单个预测的局部解释,本文提出一种新框架与交互式工具,将概念解释拓展至机制可解释性领域。该方法通过分析高层语义属性(称为概念)如何在模型内部组件中生成、交互并传播,实现对模型行为的全局剖析。不同于以往孤立分析单个神经元或预测的做法,本框架系统量化了语义概念在各层的表征方式,揭示支撑模型决策的潜在电路与信息流。核心创新是名为BAGEL(Bias Analysis with a Graph for global Explanation Layers)的可视化平台,以结构化知识图谱形式呈现分析结果,支持用户探索概念与类别间的关系,识别虚假相关性,并提升模型可信度。框架具备模型无关性与可扩展性,有助于深入理解深度学习模型在数据集偏差下的泛化能力(或失效原因)。演示地址:https://knowledge-graph-ui-4a7cb5.gitlab.io/。
原文摘要 · Abstract (English)
While concept-based interpretability methods have traditionally focused on local explanations of neural network predictions, we propose a novel framework and interactive tool that extends these methods into the domain of mechanistic interpretability. Our approach enables a global dissection of model behavior by analyzing how high-level semantic attributes (referred to as concepts) emerge, interact, and propagate through internal model components. Unlike prior work that isolates individual neurons or predictions, our framework systematically quantifies how semantic concepts are represented across layers, revealing latent circuits and information flow that underlie model decision-making. A key innovation is our visualization platform that we named BAGEL (for Bias Analysis with a Graph for global Explanation Layers), which presents these insights in a structured knowledge graph, allowing users to explore concept-class relationships, identify spurious correlations, and enhance model trustworthiness. Our framework is model-agnostic, scalable, and contributes to a deeper understanding of how deep learning models generalize (or fail to) in the presence of dataset biases. The demonstration is available at https://knowledge-graph-ui-4a7cb5.gitlab.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。