arXiv:2603.17717cs.CRcs.AI2026-03中稿 · IEEE ISCC 2026, Po…

用多模态数据训练稳定检测模型,生成高保真合成攻击数据用于评估。

Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation

  • 构建统一特征空间的多源网络攻击数据集,融合流量、包载荷与时间上下文。
  • 机器学习模型在跨数据集上实现高稳定性检测,合成数据保持高真实度与实用性。
  • 适合安全研究者与模型评估人员,提供可量化的生成数据质量评测框架。

监督式网络攻击检测是网络入侵检测系统(NIDS)的核心环节。面对当前人工智能时代下日益复杂的攻击手段,如生成式AI和强化学习的应用,保护分散于网络中的个人数据变得尤为关键。本文针对两个任务展开研究:首先,在首个统一多模态的NIDS数据集上,整合了重新处理后的CIC-IDS-2017、CIC-IoT-2023、UNSW-NB15和CIC-DDoS-2019数据集,采用相同特征空间,利用机器学习算法结合分层交叉验证,实现稳定可靠的攻击检测。其次,通过对抗学习生成合成数据,使用SDV框架、f散度、可区分性及非参数统计检验对合成数据的真实性、实用性和隐私性进行评估。结果表明,结合Synthetic Data Vault框架、TRTS与TSTR测试以及非参数统计方法,可构建具有高保真度与高实用性的生成模型,为下一代安全检测提供可靠数据支撑。

原文摘要 · Abstract (English)

Supervised detection of network attacks has always been a critical part of network intrusion detection systems (NIDS). Nowadays, in a pivotal time for artificial intelligence (AI), with even more sophisticated attacks that utilize advanced techniques, such as generative artificial intelligence (GenAI) and reinforcement learning, it has become a vital component if we wish to protect our personal data, which are scattered across the web. In this paper, we address two tasks, in the first unified multi-modal NIDS dataset, which incorporates flow-level data, packet payload information and temporal contextual features, from the reprocessed CIC-IDS-2017, CIC-IoT-2023, UNSW-NB15 and CIC-DDoS-2019, with the same feature space. In the first task we use machine learning (ML) algorithms, with stratified cross validation, in order to prevent network attacks, with stability and reliability. In the second task we use adversarial learning algorithms to generate synthetic data, compare them with the real ones and evaluate their fidelity, utility and privacy using the SDV framework, f-divergences, distinguishability and non-parametric statistical tests. The findings provide stable ML models for intrusion detection and generative models with high fidelity and utility, by combining the Synthetic Data Vault framework, the TRTS and TSTR tests, with non-parametric statistical tests and f-divergence measures.

网络入侵检测生成对抗模型数据合成安全性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。