arXiv:2512.03375cs.LG2025-12被引 3

用多模态生成模型提升入侵检测的数据平衡与准确性

MAGE-ID: A Multimodal Generative Framework for Intrusion Detection Systems

  • 融合表格数据与图像特征的扩散生成框架
  • 在两个数据集上检测准确率显著优于现有方法
  • 适合安全研究者和需要数据增强的系统开发者

现代入侵检测系统面临异构网络流量、不断演变的网络威胁以及良性与攻击流量间显著的数据不平衡问题。尽管生成模型在数据增强方面展现出潜力,但现有方法仅限于单一模态,难以捕捉跨域依赖关系。本文提出 MAGE-ID(多模态攻击生成器用于入侵检测),一种基于扩散模型的生成框架,通过统一潜在先验将表格流特征与其转换后的图像联合建模。采用基于 Transformer 与 CNN 的变分编码器,结合 EDM 风格去噪器进行联合训练,实现均衡且一致的多模态合成。在 CIC-IDS-2017 与 NSL-KDD 数据集上的评估表明,相较于 TabSyn 与 TabDDPM,MAGE-ID 在保真度、多样性及下游检测性能上均有显著提升,验证了其在多模态入侵检测数据增强中的有效性。

原文摘要 · Abstract (English)

Modern Intrusion Detection Systems (IDS) face severe challenges due to heterogeneous network traffic, evolving cyber threats, and pronounced data imbalance between benign and attack flows. While generative models have shown promise in data augmentation, existing approaches are limited to single modalities and fail to capture cross-domain dependencies. This paper introduces MAGE-ID (Multimodal Attack Generator for Intrusion Detection), a diffusion-based generative framework that couples tabular flow features with their transformed images through a unified latent prior. By jointly training Transformer and CNN-based variational encoders with an EDM style denoiser, MAGE-ID achieves balanced and coherent multimodal synthesis. Evaluations on CIC-IDS-2017 and NSL-KDD demonstrate significant improvements in fidelity, diversity, and downstream detection performance over TabSyn and TabDDPM, highlighting the effectiveness of MAGE-ID for multimodal IDS augmentation.

入侵检测生成模型多模态数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。