arXiv:2409.20427stat.MLcs.AI2024-09被引 5

提出特征重要性新框架,突破传统解释方法的局限。

Sufficient and Necessary Explanations (and What Lies in Between)

  • 构建必要性与充分性统一的连续谱解释框架
  • 发现传统方法会遗漏关键特征,新框架可补足缺陷
  • 适用于高风险决策场景的模型可解释性分析

随着复杂机器学习模型在高风险决策中的广泛应用,理解其预测机制至关重要。后验解释方法通过识别输入中对模型输出影响显著的特征提供洞察。本文形式化并研究了两类通用机器学习模型的特征重要性定义:必要性与充分性。我们表明,尽管这两类解释直观且简洁,却可能无法完整揭示模型所依赖的关键特征。为此,我们提出一种统一的重要性概念,通过探索必要性-充分性轴上的连续谱来克服上述局限。研究表明,该统一框架与基于条件独立性、博弈论(如Shapley值)等流行定义存在紧密联系。更重要的是,统一视角使我们能够识别出单一方法可能遗漏的重要特征。

原文摘要 · Abstract (English)

As complex machine learning models continue to find applications in high-stakes decision-making scenarios, it is crucial that we can explain and understand their predictions. Post-hoc explanation methods provide useful insights by identifying important features in an input $\mathbf{x}$ with respect to the model output $f(\mathbf{x})$. In this work, we formalize and study two precise notions of feature importance for general machine learning models: sufficiency and necessity. We demonstrate how these two types of explanations, albeit intuitive and simple, can fall short in providing a complete picture of which features a model finds important. To this end, we propose a unified notion of importance that circumvents these limitations by exploring a continuum along a necessity-sufficiency axis. Our unified notion, we show, has strong ties to other popular definitions of feature importance, like those based on conditional independence and game-theoretic quantities like Shapley values. Crucially, we demonstrate how a unified perspective allows us to detect important features that could be missed by either of the previous approaches alone.

可解释性特征重要性机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。