XAI解释器可能被攻击,导致安全模型决策失效。
Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
- 测试六种攻击方法篡改SHAP、LIME等解释结果
- 在钓鱼、恶意软件等场景中成功误导模型判断
- 揭示XAI系统脆弱性,警示安全应用需加强防护
可解释人工智能(XAI)帮助研究人员剖析黑箱模型的决策过程,提升信任与透明度。然而,特定XAI方法生成的解释可能遭受对抗性攻击,进而扭曲模型输出。本文研究了公平洗白解释(FE)、操纵解释(ME)和后门触发攻击(BD)等典型攻击手段,针对SHAP、LIME和IG等后置解释方法,在钓鱼、恶意软件、入侵检测和欺诈网站识别等网络安全应用场景中,共评估六种攻击流程。实验表明这些攻击能有效干扰解释结果,影响模型决策,凸显提升XAI系统抗攻击能力的紧迫性。
原文摘要 · Abstract (English)
Explainable Artificial Intelligence (XAI) has aided machine learning (ML) researchers with the power of scrutinizing the decisions of the black-box models. XAI methods enable looking deep inside the models' behavior, eventually generating explanations along with a perceived trust and transparency. However, depending on any specific XAI method, the level of trust can vary. It is evident that XAI methods can themselves be a victim of post-adversarial attacks that manipulate the expected outcome from the explanation module. Among such attack tactics, fairwashing explanation (FE), manipulation explanation (ME), and backdoor-enabled manipulation attacks (BD) are the notable ones. In this paper, we try to understand these adversarial attack techniques, tactics, and procedures (TTPs) on explanation alteration and thus the effect on the model's decisions. We have explored a total of six different individual attack procedures on post-hoc explanation methods such as SHAP (SHapley Additive exPlanations), LIME (Local Interpretable Model-agnostic Explanation), and IG (Integrated Gradients), and investigated those adversarial attacks in cybersecurity applications scenarios such as phishing, malware, intrusion, and fraudulent website detection. Our experimental study reveals the actual effectiveness of these attacks, thus providing an urgency for immediate attention to enhance the resiliency of XAI methods and their applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。