让深度模型能回答'如果...会怎样'的问题,提升可解释性。
From Black-box to Causal-box: Towards Building More Interpretable Models
- 提出因果可解释性框架,判断模型能否回答反事实问题。
- 发现可解释性与预测精度存在根本权衡,需选择最优特征集。
- 适用于医疗、金融等高风险场景,需理解模型推理过程的用户。
深度学习模型的预测理解仍是重大挑战,尤其在高风险应用中。一种有前景的方法是让模型具备回答反事实问题的能力——即超越观测数据的假设性‘如果…会怎样’问题,从而揭示模型推理逻辑。本文引入因果可解释性概念,形式化了在何种条件下特定模型和观测数据可支持反事实查询。分析两类常见模型——黑箱模型与基于概念的预测器——表明二者通常不具备因果可解释性。为弥补这一缺陷,我们构建了一个从设计上实现因果可解释性的框架。具体而言,推导出一个完整的图结构判据,用于判断给定模型架构是否支持特定反事实查询。该框架揭示了因果可解释性与预测准确率之间的根本权衡,并通过识别唯一最大特征集,实现可解释性与预测表达力的最优平衡。实验验证了理论结果。
原文摘要 · Abstract (English)
Understanding the predictions made by deep learning models remains a central challenge, especially in high-stakes applications. A promising approach is to equip models with the ability to answer counterfactual questions -- hypothetical ``what if?'' scenarios that go beyond the observed data and provide insight into a model reasoning. In this work, we introduce the notion of causal interpretability, which formalizes when counterfactual queries can be evaluated from a specific class of models and observational data. We analyze two common model classes -- blackbox and concept-based predictors -- and show that neither is causally interpretable in general. To address this gap, we develop a framework for building models that are causally interpretable by design. Specifically, we derive a complete graphical criterion that determines whether a given model architecture supports a given counterfactual query. This leads to a fundamental tradeoff between causal interpretability and predictive accuracy, which we characterize by identifying the unique maximal set of features that yields an interpretable model with maximal predictive expressiveness. Experiments corroborate the theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。