arXiv:2603.08901cs.CRcs.AI2026-03被引 3

用扩散模型生成难辨别的恶意网络流量,骗过深度学习检测系统

NetDiffuser: Deceiving DNN-Based Network Attack Detection Systems with Diffusion-Generated Adversarial Traffic

  • 基于特征分类与扩散模型,生成语义一致的伪装流量
  • 攻击成功率提升29.93%,检测器AUC-ROC下降最多0.534
  • 适合研究防御机制或测试系统安全性的研究人员

基于深度学习的网络入侵检测系统(NIDS)在识别恶意网络流量方面展现出巨大潜力,但其易受对抗样本(AEs)攻击。现有攻击多通过扰动数据制造误分类,而自然对抗样本(NAEs)因与真实数据高度相似,更难被发现。本文提出NetDiffuser,一种生成NAEs的新框架,包含两项创新:一是设计特征分类算法,识别相对独立的网络特征,扰动时保持流量有效性;二是首次将扩散模型应用于注入语义一致的扰动,生成逼真对抗流量。在三个基准NIDS数据集上,针对多种模型架构和先进检测器的实验表明,NetDiffuser攻击成功率最高提升29.93%,且在部分情况下使检测器的AUC-ROC得分降低至少0.267(最高达0.534)。

原文摘要 · Abstract (English)

Deep learning (DL)-based Network Intrusion Detection System (NIDS) has demonstrated great promise in detecting malicious network traffic. However, they face significant security risks due to their vulnerability to adversarial examples (AEs). Most existing adversarial attacks maliciously perturb data to maximize misclassification errors. Among AEs, natural adversarial examples (NAEs) are particularly difficult to detect because they closely resemble real data, making them challenging for both humans and machine learning models to distinguish from legitimate inputs. Creating NAEs is crucial for testing and strengthening NIDS defenses. This paper proposes NetDiffuser1, a novel framework for generating NAEs capable of deceiving NIDS. NetDiffuser consists of two novel components. First, a new feature categorization algorithm is designed to identify relatively independent features in network traffic. Perturbing these features minimizes changes while preserving network flow validity. The second component is a novel application of diffusion models to inject semantically consistent perturbations for generating NAEs. NetDiffuser performance was extensively evaluated using three benchmark NIDS datasets across various model architectures and state-of-the-art adversarial detectors. Our experimental results show that NetDiffuser achieves up to a 29.93% higher attack success rate and reduces AE detection performance by at least 0.267 (in some cases up to 0.534) in the Area under the Receiver Operating Characteristic Curve (AUC-ROC) score compared to the baseline attacks.

对抗攻击网络检测扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。