arXiv:2511.15371cs.LG2025-11中稿 · Northern Lights De…被引 1

通过反事实分布量化特征重要性,提升模型解释的准确性。

CID: Measuring Feature Importance Through Counterfactual Distributions

  • 构建正负反事实样本,用核密度估计其分布
  • 基于分布差异度量排序特征重要性,性能优于现有方法
  • 适合需要高可信度解释的模型分析场景

机器学习中评估单个特征的重要性对于理解模型决策过程至关重要。尽管已有多种方法,但缺乏明确的比较基准,亟需更可靠的度量方式。本文提出一种新的后处理局部特征重要性方法——反事实重要性分布(CID)。通过生成正负两类反事实样本,利用核密度估计建模其分布,并基于分布差异度量对特征进行排序。该度量具有严格的数学基础,满足作为有效度量所需的关键性质。实验表明,相比主流局部解释方法,CID不仅提供互补视角,还在忠实性指标(充分性与完备性)上表现更优,生成更可信的解释结果。这凸显了其在模型分析中的潜在价值。

原文摘要 · Abstract (English)

Assessing the importance of individual features in Machine Learning is critical to understand the model's decision-making process. While numerous methods exist, the lack of a definitive ground truth for comparison highlights the need for alternative, well-founded measures. This paper introduces a novel post-hoc local feature importance method called Counterfactual Importance Distribution (CID). We generate two sets of positive and negative counterfactuals, model their distributions using Kernel Density Estimation, and rank features based on a distributional dissimilarity measure. This measure, grounded in a rigorous mathematical framework, satisfies key properties required to function as a valid metric. We showcase the effectiveness of our method by comparing with well-established local feature importance explainers. Our method not only offers complementary perspectives to existing approaches, but also improves performance on faithfulness metrics (both for comprehensiveness and sufficiency), resulting in more faithful explanations of the system. These results highlight its potential as a valuable tool for model analysis.

特征重要性模型解释反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。