arXiv:2601.10587cs.CVcs.AI2026-01

用SHAP值生成肉眼难辨的视觉欺骗攻击,让模型误判。

Adversarial Evasion Attacks on Computer Vision using SHAP Values

  • 基于SHAP值量化输入对模型输出的影响,设计白盒攻击
  • 在梯度隐藏场景下,比FGSM方法更易引发模型误判
  • 攻击几乎不可见,适合研究模型鲁棒性或安全防御

本文提出一种基于SHAP值的白盒对抗逃避攻击,用于计算机视觉模型。该攻击通过量化推理阶段各输入对输出的重要性,在保持人类感知不可察觉的前提下,降低模型输出置信度或诱导误分类。由于其对人眼几乎不可见,此类攻击极具隐蔽性。实验对比了SHAP攻击与经典的快速梯度符号法(Fast Gradient Sign Method, FGSM),结果表明在梯度隐藏场景下,SHAP攻击能更稳定地引发模型误判,表现出更强的鲁棒性。

原文摘要 · Abstract (English)

The paper introduces a white-box attack on computer vision models using SHAP values. It demonstrates how adversarial evasion attacks can compromise the performance of deep learning models by reducing output confidence or inducing misclassifications. Such attacks are particularly insidious as they can deceive the perception of an algorithm while eluding human perception due to their imperceptibility to the human eye. The proposed attack leverages SHAP values to quantify the significance of individual inputs to the output at the inference stage. A comparison is drawn between the SHAP attack and the well-known Fast Gradient Sign Method. We find evidence that SHAP attacks are more robust in generating misclassifications particularly in gradient hiding scenarios.

对抗攻击SHAP值模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。