arXiv:2409.14590cs.LGcs.AI2024-09被引 10

现有可解释AI方法失效,因未正确定义问题与评估标准。

Explainable AI needs formalization

  • 明确问题定义,构建可验证的解释框架
  • 当前方法对无关特征错误赋权,导致结果不可靠
  • 推动针对不同场景的解释正确性标准建立

可解释人工智能(XAI)旨在使机器学习决策具备人类可理解性,但现有方法存在根本缺陷:无法可靠回答关于模型、训练数据或测试输入的相关问题,因其系统性地将重要性归因于与预测目标无关的输入特征。这限制了XAI在诊断数据与模型、科学发现及干预目标识别中的应用价值。根源在于当前方法未解决明确问题,也未以针对性解释正确性标准进行评估。研究者应正式定义所要解决的问题,并据此设计方法,从而发展出依赖具体应用场景的解释正确性概念与客观性能度量,用于验证XAI算法的有效性。

原文摘要 · Abstract (English)

The field of "explainable artificial intelligence" (XAI) seemingly addresses the desire that decisions of machine learning systems should be human-understandable. However, in its current state, XAI itself needs scrutiny. Popular methods cannot reliably answer relevant questions about ML models, their training data, or test inputs, because they systematically attribute importance to input features that are independent of the prediction target. This limits the utility of XAI for diagnosing and correcting data and models, for scientific discovery, and for identifying intervention targets. The fundamental reason for this is that current XAI methods do not address well-defined problems and are not evaluated against targeted criteria of explanation correctness. Researchers should formally define the problems they intend to solve and design methods accordingly. This will lead to diverse use-case-dependent notions of explanation correctness and objective metrics of explanation performance that can be used to validate XAI algorithms.

可解释AI形式化评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。