arXiv:2410.15648cs.LGstat.ME2024-10被引 3

揭示模型干预效应何时具有因果解释力,提升模型可信度。

Linking Model Intervention to Causal Interpretation in Model Explanation

  • 通过数学条件判断特征干预是否具因果意义。
  • 在存在未观测变量时,干预效应可能无法反映真实因果。
  • 适用于需验证模型决策可靠性的领域专家。

模型解释中常使用干预直觉:通过改变特征值从当前值到基线值,量化其对模型输出的影响。然而,这种干预效应本质上是关联性而非因果性。本文研究了在何种条件下,直观的模型干预效应可具备因果解释能力,即能否表明某特征是结果的直接原因。该工作将模型干预与因果解释相连接,这对判断机器学习模型是否值得领域专家信任至关重要。研究揭示了在存在未观测特征的环境下,使用干预效应进行因果解释的局限性。通过半合成数据集实验验证了相关定理,并展示了干预效应在模型解释中的潜在应用价值。

原文摘要 · Abstract (English)

Intervention intuition is often used in model explanation where the intervention effect of a feature on the outcome is quantified by the difference of a model prediction when the feature value is changed from the current value to the baseline value. Such a model intervention effect of a feature is inherently association. In this paper, we will study the conditions when an intuitive model intervention effect has a causal interpretation, i.e., when it indicates whether a feature is a direct cause of the outcome. This work links the model intervention effect to the causal interpretation of a model. Such an interpretation capability is important since it indicates whether a machine learning model is trustworthy to domain experts. The conditions also reveal the limitations of using a model intervention effect for causal interpretation in an environment with unobserved features. Experiments on semi-synthetic datasets have been conducted to validate theorems and show the potential for using the model intervention effect for model interpretation.

模型解释因果推断可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。