arXiv:2412.05592cs.AI2024-12ECCV被引 7

XAI评估中的超参数灵活性可被操纵,影响结果可靠性。

From Flexibility to Manipulation: The Slippery Slope of XAI Evaluation

  • 将超参数设置视为对抗攻击,改变可显著影响评估结果
  • 多数据集测试显示不同方法与模型的评估结果变化巨大
  • 提出基于排名的策略以增强评估对操纵的鲁棒性

可解释人工智能(XAI)的定量评估面临缺乏真实解释标签的根本挑战。当评估方法包含大量需用户指定的超参数时,这一问题尤为突出,因为没有真实标签可判断最优超参数选择。由于无法进行超参数的穷举搜索,研究者通常参照文献中的惯例进行选择,这为用户提供了极大的灵活性。本文揭示了这种灵活性如何被用来操纵评估结果。我们将这种操纵视为对评估过程的对抗攻击:看似无害的超参数调整会显著改变评估结果。我们在多个数据集上验证了该操纵的有效性,发现不同解释方法和模型的评估结果出现显著波动。最后,我们提出一种基于超参数排序的缓解策略,旨在提升评估对操纵的鲁棒性。本工作凸显了实现可靠XAI评估的困难,并强调评估过程需要整体化与透明化。

原文摘要 · Abstract (English)

The lack of ground truth explanation labels is a fundamental challenge for quantitative evaluation in explainable artificial intelligence (XAI). This challenge becomes especially problematic when evaluation methods have numerous hyperparameters that must be specified by the user, as there is no ground truth to determine an optimal hyperparameter selection. It is typically not feasible to do an exhaustive search of hyperparameters so researchers typically make a normative choice based on similar studies in the literature, which provides great flexibility for the user. In this work, we illustrate how this flexibility can be exploited to manipulate the evaluation outcome. We frame this manipulation as an adversarial attack on the evaluation where seemingly innocent changes in hyperparameter setting significantly influence the evaluation outcome. We demonstrate the effectiveness of our manipulation across several datasets with large changes in evaluation outcomes across several explanation methods and models. Lastly, we propose a mitigation strategy based on ranking across hyperparameters that aims to provide robustness towards such manipulation. This work highlights the difficulty of conducting reliable XAI evaluation and emphasizes the importance of a holistic and transparent approach to evaluation in XAI.

XAI评估可解释性对抗攻击超参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。