arXiv:2603.24524cs.LGcs.AI2026-03中稿 · the Fourth World C…

提出多维度评估框架,揭示单一指标无法全面衡量不确定性归因效果。

No Single Metric Tells the Whole Story: A Multi-Dimensional Evaluation Framework for Uncertainty Attributions

  • 基于XAI的四个核心属性构建评估框架,新增可传达性指标
  • 梯度法在一致性和可传达性上优于扰动法,蒙特卡洛丢弃表现更优
  • 强调需多指标联合评估,适合方法开发者与评测研究者

可解释人工智能(XAI)研究常聚焦于模型预测的解释,近年有方法尝试通过将不确定性归因于输入特征来解释预测不确定性。然而,现有评估方法依赖异质性代理任务和指标,导致结果难以比较。本文将不确定性归因与成熟的Co-12 XAI评估框架对齐,提出正确性、一致性、连续性和紧凑性四类属性的具体实现,并引入专为不确定性归因设计的可传达性属性,评估当认知不确定性增加时,特征级归因是否可靠传递。我们在表格和图像数据上,使用八种不同不确定性量化与特征归因方法组合进行实验。结果表明,梯度法在一致性和可传达性上持续优于扰动法,蒙特卡洛丢弃优于蒙特卡洛丢弃。尽管多数指标对方法排序一致,但方法间一致性仍较低。这说明单一指标无法充分评估不确定性归因质量。该框架为系统性比较与开发不确定性归因方法奠定了基础。

原文摘要 · Abstract (English)

Research on explainable AI (XAI) has frequently focused on explaining model predictions. More recently, methods have been proposed to explain prediction uncertainty by attributing it to input features (uncertainty attributions). However, the evaluation of these methods remains inconsistent as studies rely on heterogeneous proxy tasks and metrics, hindering comparability. We address this by aligning uncertainty attributions with the well-established Co-12 framework for XAI evaluation. We propose concrete implementations for the correctness, consistency, continuity, and compactness properties. Additionally, we introduce conveyance, a property tailored to uncertainty attributions that evaluates whether controlled increases in epistemic uncertainty reliably propagate to feature-level attributions. We demonstrate our evaluation framework with eight metrics across combinations of uncertainty quantification and feature attribution methods on tabular and image data. Our experiments show that gradient-based methods consistently outperform perturbation-based approaches in consistency and conveyance, while Monte-Carlo dropconnect outperforms Monte-Carlo dropout in most metrics. Although most metrics rank the methods consistently across samples, inter-method agreement remains low. This suggests no single metric sufficiently evaluates uncertainty attribution quality. The proposed evaluation framework contributes to the body of knowledge by establishing a foundation for systematic comparison and development of uncertainty attribution methods.

可解释AI不确定性归因评估框架多指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。