arXiv:2411.09639cs.LG2024-11

解决概念缺失时的因果解释偏差,提升模型可解释性。

MCCE: Missingness-aware Causal Concept Explainer

  • 引入新框架MCCE,自动校正缺失概念带来的偏差
  • 在真实数据集上优于现有方法,效果稳定可靠
  • 适合需要处理不完整标注数据的可解释性研究

因果概念效应估计在可解释机器学习领域日益受到关注。该方法通过估计人类可理解概念对模型行为的因果效应来解释模型,这些概念比原始输入(如词元)更易理解。然而,现有方法假设数据集中所有概念均被完整观测,这在实际中常因标注不全或概念缺失而失效。我们理论证明,未观测概念会偏倚对可观测概念因果效应的估计。为此,我们提出缺失感知因果概念解释器(MCCE),一种专为概念不完全可观测场景设计的新框架。该框架学习残余偏差并利用线性预测器建模概念与黑箱模型输出之间的关系,支持局部与全局解释。在真实数据集上的验证表明,MCCE在因果概念效应估计方面显著优于当前最优方法。

原文摘要 · Abstract (English)

Causal concept effect estimation is gaining increasing interest in the field of interpretable machine learning. This general approach explains the behaviors of machine learning models by estimating the causal effect of human-understandable concepts, which represent high-level knowledge more comprehensibly than raw inputs like tokens. However, existing causal concept effect explanation methods assume complete observation of all concepts involved within the dataset, which can fail in practice due to incomplete annotations or missing concept data. We theoretically demonstrate that unobserved concepts can bias the estimation of the causal effects of observed concepts. To address this limitation, we introduce the Missingness-aware Causal Concept Explainer (MCCE), a novel framework specifically designed to estimate causal concept effects when not all concepts are observable. Our framework learns to account for residual bias resulting from missing concepts and utilizes a linear predictor to model the relationships between these concepts and the outputs of black-box machine learning models. It can offer explanations on both local and global levels. We conduct validations using a real-world dataset, demonstrating that MCCE achieves promising performance compared to state-of-the-art explanation methods in causal concept effect estimation.

可解释性因果推断缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。