arXiv:2411.02470cs.CVcs.AI2024-11AAAI被引 12

用真人评价标准评估AI解释,让解释更贴近人类认知。

Benchmarking XAI Explanations with Human-Aligned Evaluations

  • 构建首个涵盖多种模型和解释方法的大规模人类评价数据集
  • 通过数据驱动方法预测人对解释的偏好,评分可靠可扩展
  • 支持跨模态解释对比,适合研究可解释AI的学者

我们提出PASTA(Perceptual Assessment System for explanation of Artificial Intelligence),一种以人为中心的计算机视觉可解释AI(XAI)评估框架。首个贡献是创建PASTA-dataset,首个大规模基准数据集,覆盖多样化的模型及基于显著性与概念的解释方法,支持基于人类判断的稳健、可比分析。第二个贡献是基于该数据集的自动化、数据驱动评估方法,称为PASTA-score,能可靠预测人类偏好,实现可扩展、一致的评价。此外,本基准首次支持跨模态解释的比较。我们进一步提出将此评分方法用于探测现有模型的可解释性,并构建更符合人类认知的XAI方法。

原文摘要 · Abstract (English)

We introduce PASTA (Perceptual Assessment System for explanaTion of Artificial Intelligence), a novel human-centric framework for evaluating eXplainable AI (XAI) techniques in computer vision. Our first contribution is the creation of the PASTA-dataset, the first large-scale benchmark that spans a diverse set of models and both saliency-based and concept-based explanation methods. This dataset enables robust, comparative analysis of XAI techniques based on human judgment. Our second contribution is an automated, data-driven benchmark that predicts human preferences using the PASTA-dataset. This scoring called PASTA-score method offers scalable, reliable, and consistent evaluation aligned with human perception. Additionally, our benchmark allows for comparisons between explanations across different modalities, an aspect previously unaddressed. We then propose to apply our scoring method to probe the interpretability of existing models and to build more human interpretable XAI methods.

可解释AI人类评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。