用人工生成的解释真值评估XAI方法,解决无标准答案的难题。
Exploring SAIG Methods for an Objective Evaluation of XAI
- 通过构造人工真值(SAIG)实现对XAI解释的直接评估。
- 发现当前缺乏统一有效的评估方法,结果差异大。
- 适合关注XAI评测标准的研究者和从业者。
可解释人工智能(XAI)的评估是快速发展的领域,方法多样且复杂。与传统AI评估不同,解释本身没有唯一正确答案,导致客观评估困难。本文提出一种新方向:使用合成人工智能真值(SAIG)方法生成人工解释真值,从而直接评估XAI技术。这是首次对SAIG方法的系统性综述与分析。我们提出一个全新分类体系,识别出七项关键特征以区分不同方法。对比研究揭示了当前在最优评估技术上缺乏共识,凸显该领域亟需进一步研究与标准统一。
原文摘要 · Abstract (English)
The evaluation of eXplainable Artificial Intelligence (XAI) methods is a rapidly growing field, characterized by a wide variety of approaches. This diversity highlights the complexity of the XAI evaluation, which, unlike traditional AI assessment, lacks a universally correct ground truth for the explanation, making objective evaluation challenging. One promising direction to address this issue involves the use of what we term Synthetic Artificial Intelligence Ground truth (SAIG) methods, which generate artificial ground truths to enable the direct evaluation of XAI techniques. This paper presents the first review and analysis of SAIG methods. We introduce a novel taxonomy to classify these approaches, identifying seven key features that distinguish different SAIG methods. Our comparative study reveals a concerning lack of consensus on the most effective XAI evaluation techniques, underscoring the need for further research and standardization in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。