arXiv:2410.15491cs.LGstat.ME2024-10

用因果模型动态发现任务相关概念,提升模型可解释性。

Structural Causality-based Generalizable Concept Discovery Models

  • 基于变分自编码器学习数据生成因子,再通过因果模型构建任务特定概念。
  • 在D-sprites和Shapes3D上成功发现由生成因子驱动的任务相关概念。
  • 方法可扩展至任意数量概念,适用于任意下游任务,通用性强。

随着对可解释深度神经网络架构的需求增长,语义概念被用作可解释单元。现有方法通过解耦表示学习估计生成因子,并将其作为解释DNN的语义概念。然而,尽管数据集的生成因子固定,概念却随下游任务变化而变化。本文提出一种基于变分自编码器(VAE)的解耦机制,先学习数据的相互独立生成因子,再利用结构因果模型(SCM)学习任务特定概念。方法假设生成因子与概念构成二分图,从生成因子到概念存在有向因果边。在生成因子已知的数据集D-sprites和Shapes3D上进行实验,结果表明所提方法能有效学习出由生成因子解释的任务特定概念。与现有因果概念发现方法不同,本方法可推广至任意数量的概念,且对任意下游任务具有灵活性。

原文摘要 · Abstract (English)

The rising need for explainable deep neural network architectures has utilized semantic concepts as explainable units. Several approaches utilizing disentangled representation learning estimate the generative factors and utilize them as concepts for explaining DNNs. However, even though the generative factors for a dataset remain fixed, concepts are not fixed entities and vary based on downstream tasks. In this paper, we propose a disentanglement mechanism utilizing a variational autoencoder (VAE) for learning mutually independent generative factors for a given dataset and subsequently learning task-specific concepts using a structural causal model (SCM). Our method assumes generative factors and concepts to form a bipartite graph, with directed causal edges from generative factors to concepts. Experiments are conducted on datasets with known generative factors: D-sprites and Shapes3D. On specific downstream tasks, our proposed method successfully learns task-specific concepts which are explained well by the causal edges from the generative factors. Lastly, separate from current causal concept discovery methods, our methodology is generalizable to an arbitrary number of concepts and flexible to any downstream tasks.

可解释性因果发现概念学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。