arXiv:2511.10088cs.LGcs.AI2025-11

让模型解释变得不可信:无声篡改图像解释却不被察觉

eXIAA: eXplainable Injections for Adversarial Attack

  • 仅需预测结果和解释,单步完成黑箱攻击
  • 解释差异显著,但预测不变且人眼无法察觉
  • 揭示现有可解释性方法在关键场景中的脆弱性

后验可解释性方法旨在阐明模型行为的原因。本文提出一种新型黑箱、模型无关的对抗攻击,针对图像领域的后验可解释人工智能(XAI)。攻击目标是在不改变预测类别且人眼难以察觉的前提下,篡改原始解释。与以往方法不同,本方法无需访问模型或其权重,仅需模型输出的预测结果和解释。实验表明,该攻击可在单步内显著改变解释内容,同时保持预测概率不变。我们在ImageNet上对预训练ResNet-18和ViT-B16使用saliency maps、Integrated Gradients和DeepLIFT SHAP等方法生成解释,并系统构造攻击。结果表明,攻击导致解释发生巨大变化,但原始图像与扰动图像间的结构相似度(SSIM)仍接近1.0,证明视觉不可察觉;解释变化程度以平均绝对差衡量,显著提升。

原文摘要 · Abstract (English)

Post-hoc explainability methods are a subset of Machine Learning (ML) that aim to provide a reason for why a model behaves in a certain way. In this paper, we show a new black-box model-agnostic adversarial attack for post-hoc explainable Artificial Intelligence (XAI), particularly in the image domain. The goal of the attack is to modify the original explanations while being undetected by the human eye and maintain the same predicted class. In contrast to previous methods, we do not require any access to the model or its weights, but only to the model's computed predictions and explanations. Additionally, the attack is accomplished in a single step while significantly changing the provided explanations, as demonstrated by empirical evaluation. The low requirements of our method expose a critical vulnerability in current explainability methods, raising concerns about their reliability in safety-critical applications. We systematically generate attacks based on the explanations generated by post-hoc explainability methods (saliency maps, integrated gradients, and DeepLIFT SHAP) for pretrained ResNet-18 and ViT-B16 on ImageNet. The results show that our attacks could lead to dramatically different explanations without changing the predictive probabilities. We validate the effectiveness of our attack, compute the induced change based on the explanation with mean absolute difference, and verify the closeness of the original image and the corrupted one with the Structural Similarity Index Measure (SSIM).

对抗攻击可解释性黑箱攻击图像安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。