提出新方法量化神经网络中输入特征的内在因果贡献,解释更可信。
On Measuring Intrinsic Causal Attributions in Deep Neural Networks
- 将神经网络视为结构因果模型,定义内在因果贡献(ICC)
- 实验表明ICC比现有全局解释方法更直观可靠
- 与Sobol'指数相关,适合需要可解释性的研究者
量化神经网络中输入特征的因果影响已成为日益关注的话题。现有方法通常评估直接、间接和总因果效应。本文将神经网络视为结构因果模型(SCMs),并引入内在因果贡献(ICC)概念。我们提出一种可识别的生成式后验框架来量化ICC。同时建立了ICC与Sobol'指数之间的联系。在合成数据和真实数据集上的实验表明,相比现有全局解释技术,ICC能生成更直观、更可靠的解释。
原文摘要 · Abstract (English)
Quantifying the causal influence of input features within neural networks has become a topic of increasing interest. Existing approaches typically assess direct, indirect, and total causal effects. This work treats NNs as structural causal models (SCMs) and extends our focus to include intrinsic causal contributions (ICC). We propose an identifiable generative post-hoc framework for quantifying ICC. We also draw a relationship between ICC and Sobol' indices. Our experiments on synthetic and real-world datasets demonstrate that ICC generates more intuitive and reliable explanations compared to existing global explanation techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。