构建大规模可解释AI评估基准,揭示现有方法排名矛盾问题。
Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
- 设计包含7560种组合的大规模评测框架,覆盖17种方法与20个指标。
- 发现不同评估指标结果常冲突,导致方法排名不可靠。
- 公开326,000张显著图和378,000个评分数据,助力未来研究。
可解释人工智能(XAI)领域发展迅速,涌现出大量方法与评估指标。然而,现有研究多局限于少数方法,忽略模型架构与输入数据等关键设计参数的影响,且仅依赖一两个指标,缺乏充分验证,易引发选择偏差并忽视指标间的差异。为解决此问题,本文提出LATEC——一个大规模基准,系统评估17种主流XAI方法,使用20个不同指标,在7,560种组合条件下进行分析。结果显示,多个指标间存在显著冲突,导致方法排名不可靠,因此提出更稳健的评估方案。此外,我们全面评估各类方法以辅助实践者选型。值得注意的是,表现最佳的新方法‘期望梯度’此前未被相关研究涵盖。LATEC通过公开全部326,000张显著图和378,000个指标得分,建立了一个(元)评估数据集。基准代码与数据已开源:https://github.com/IML-DKFZ/latec。
原文摘要 · Abstract (English)
Explainable AI (XAI) is a rapidly growing domain with a myriad of proposed methods as well as metrics aiming to evaluate their efficacy. However, current studies are often of limited scope, examining only a handful of XAI methods and ignoring underlying design parameters for performance, such as the model architecture or the nature of input data. Moreover, they often rely on one or a few metrics and neglect thorough validation, increasing the risk of selection bias and ignoring discrepancies among metrics. These shortcomings leave practitioners confused about which method to choose for their problem. In response, we introduce LATEC, a large-scale benchmark that critically evaluates 17 prominent XAI methods using 20 distinct metrics. We systematically incorporate vital design parameters like varied architectures and diverse input modalities, resulting in 7,560 examined combinations. Through LATEC, we showcase the high risk of conflicting metrics leading to unreliable rankings and consequently propose a more robust evaluation scheme. Further, we comprehensively evaluate various XAI methods to assist practitioners in selecting appropriate methods aligning with their needs. Curiously, the emerging top-performing method, Expected Gradients, is not examined in any relevant related study. LATEC reinforces its role in future XAI research by publicly releasing all 326k saliency maps and 378k metric scores as a (meta-)evaluation dataset. The benchmark is hosted at: https://github.com/IML-DKFZ/latec.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。