arXiv:2501.01209cs.AIcs.LG2025-01

用重描述框架解析深度模型,让复杂神经网络变得可解释。

A redescription mining framework for post-hoc explaining and relating deep learning models

  • 通过统计显著性挖掘神经元激活的重描述模式。
  • 可关联神经元与标签、属性,跨模型或层进行对比分析。
  • 不依赖网络结构,支持多标签等复杂场景,适合研究者使用。

深度学习模型在结构化和非结构化数据上表现优异,极大拓展了机器学习的应用范围。其在预测、模式识别和生成新数据方面的成功对科学与产业产生了深远影响。然而,由于模型规模庞大,解释性差。本文提出一种基于重描述的后验解释与关联框架,通过识别神经元激活的统计显著重描述,实现对任意深度学习模型的群体分析。该框架可将神经元与目标标签或描述属性关联,连接单个模型内的不同层,或关联多个模型。框架独立于神经网络架构,支持多标签、多目标等复杂场景,并能模拟教学式与分解式规则提取方法。相比现有可解释AI方法,该框架提供了差异化信息,显著提升模型的可解释性与透明度。

原文摘要 · Abstract (English)

Deep learning models (DLMs) achieve increasingly high performance both on structured and unstructured data. They significantly extended applicability of machine learning to various domains. Their success in making predictions, detecting patterns and generating new data made significant impact on science and industry. Despite these accomplishments, DLMs are difficult to explain because of their enormous size. In this work, we propose a novel framework for post-hoc explaining and relating DLMs using redescriptions. The framework allows cohort analysis of arbitrary DLMs by identifying statistically significant redescriptions of neuron activations. It allows coupling neurons to a set of target labels or sets of descriptive attributes, relating layers within a single DLM or associating different DLMs. The proposed framework is independent of the artificial neural network architecture and can work with more complex target labels (e.g. multi-label or multi-target scenario). Additionally, it can emulate both pedagogical and decompositional approach to rule extraction. The aforementioned properties of the proposed framework can increase explainability and interpretability of arbitrary DLMs by providing different information compared to existing explainable-AI approaches.

可解释性深度学习重描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。