arXiv:2502.17007cs.LGcs.AI2025-02被引 5

用不确定性量化统一反事实解释,让模型解释更可靠。

Uncertainty Quantification as a Principled Foundation for Explainable Artificial Intelligence: A Case Study of Counterfactual Explanations

  • 将反事实解释的核心属性转化为不确定性表达
  • 两种设计均在多个基准上达到顶尖性能
  • 适合追求可信赖AI解释的研究者与工程师

本文指出,透明性研究忽视了人工智能的诸多基础概念。以反事实可解释性中的不确定性量化为例,我们证明其广泛采用可解决该领域关键挑战。通过将核心反事实属性表述为不确定性,我们构建了两种解释器:一种仅基于不确定性估计,另一种结合特征空间距离。大量实验表明,尽管设计极为简单,本框架性能仍显著优于众多前沿方法。更广泛而言,将人工智能基础融入透明性研究,有望实现更可靠、鲁棒且易理解的预测模型。我们认为,使可解释性真正具备不确定性感知能力,是迈向这一目标的第一步。

原文摘要 · Abstract (English)

In this paper we argue that, to its detriment, transparency research overlooks many foundational concepts of artificial intelligence. As an illustrating example we focus on uncertainty quantification in the context of counterfactual explainability, demonstrating that its broader adoption could address key challenges in the field. To this end, we show how uncertainty can provide a principled unifying framework for counterfactual explainability by expressing the core counterfactual properties in terms of uncertainty, allowing us to build two variants of an explainer upon them -- one based solely on uncertainty estimates and another pairing them with distance measured in the feature space. Our comprehensive experiments illustrate highly competitive performance of our framework when compared to many state-of-the-art methods despite its radically simple design. More broadly, the paper demonstrates that integrating artificial intelligence fundamentals into transparency research promises to yield more reliable, robust and understandable predictive models. We posit that making artificial intelligence explainability truly uncertainty-aware is the first step towards this goal.

可解释AI不确定性量化反事实解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。