arXiv:2412.10942cs.CV2024-12被引 4

提出新方法评估XAI稳定性度量的可靠性,发现现有指标无法识别随机解释。

Meta-evaluating stability measures: MAX-Senstivity & AVG-Sensitivity

  • 构建元评估框架,检验稳定性度量的有效性
  • 两种指标在随机解释下均失效,无法区分好坏
  • 适合关注XAI可信性与评估标准的研究者

可解释人工智能(XAI)系统的发展带来了诸多挑战,其中鲁棒性或稳定性一直是研究重点。尽管已有多个客观评估指标被提出,但相关问题仍存。本文提出一种新型元评估方法,用于分析稳定性度量本身的正确性。我们设计了两项新测试,评估两种常用指标:AVG-Sensitivity与MAX-Sensitivity。测试场景包括决策树生成的完美与稳健解释,以及完全随机的解释与预测。结果显示,这两种指标均无法识别随机解释为错误,暴露出其整体不可靠性。

原文摘要 · Abstract (English)

The use of eXplainable Artificial Intelligence (XAI) systems has introduced a set of challenges that need resolution. The XAI robustness, or stability, has been one of the goals of the community from its beginning. Multiple authors have proposed evaluating this feature using objective evaluation measures. Nonetheless, many questions remain. With this work, we propose a novel approach to meta-evaluate these metrics, i.e. analyze the correctness of the evaluators. We propose two new tests that allowed us to evaluate two different stability measures: AVG-Sensitiviy and MAX-Senstivity. We tested their reliability in the presence of perfect and robust explanations, generated with a Decision Tree; as well as completely random explanations and prediction. The metrics results showed their incapacity of identify as erroneous the random explanations, highlighting their overall unreliability.

XAI稳定性评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。