arXiv:2602.06395cs.CRcs.AI2026-02中稿 · publication in 18t…被引 1

研究黑客如何用微小改动骗过网络安全模型,并发现模型越难被攻破,解释性就越差。

Empirical Analysis of Adversarial Robustness and Explainability Drift in Cybersecurity Classifiers

  • 用对抗攻击测试模型,量化鲁棒性与解释性变化
  • 对抗训练让模型鲁棒性提升9%,且不降低正常数据准确率
  • 适合关注AI安全可信性的研究人员和工程师

机器学习模型在网络钓鱼检测和网络入侵防御等网络安全应用中日益普及,但其仍易受对抗扰动(即精心设计的小幅输入修改)影响,导致检测准确率下降并破坏可解释性。本文对网络钓鱼网址分类和网络入侵检测两个领域进行了实证研究,评估了L∞范数约束的快速梯度符号法(FGSM)和投影梯度下降(PGD)扰动对模型准确率的影响,并引入鲁棒性指数(RI),定义为准确率-扰动曲线下的面积。基于梯度的特征敏感性和基于SHAP的归因漂移分析揭示了最易受攻击的输入特征。在Phishing Websites和UNSW NB15数据集上的实验显示,鲁棒性趋势一致:对抗训练使RI最高提升9个百分点,同时保持干净数据下的准确率。结果表明鲁棒性与可解释性退化存在耦合关系,强调了在可信赖的AI驱动网络安全系统设计中进行量化评估的重要性。

原文摘要 · Abstract (English)

Machine learning (ML) models are increasingly deployed in cybersecurity applications such as phishing detection and network intrusion prevention. However, these models remain vulnerable to adversarial perturbations small, deliberate input modifications that can degrade detection accuracy and compromise interpretability. This paper presents an empirical study of adversarial robustness and explainability drift across two cybersecurity domains phishing URL classification and network intrusion detection. We evaluate the impact of L (infinity) bounded Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD) perturbations on model accuracy and introduce a quantitative metric, the Robustness Index (RI), defined as the area under the accuracy perturbation curve. Gradient based feature sensitivity and SHAP based attribution drift analyses reveal which input features are most susceptible to adversarial manipulation. Experiments on the Phishing Websites and UNSW NB15 datasets show consistent robustness trends, with adversarial training improving RI by up to 9 percent while maintaining clean-data accuracy. These findings highlight the coupling between robustness and interpretability degradation and underscore the importance of quantitative evaluation in the design of trustworthy, AI-driven cybersecurity systems.

网络安全对抗攻击可解释性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。