生成超1200万样本的细粒度网络攻击数据集,解决旧数据标签粗、样本少问题。
Technical Report: Generating the WEB-IDS23 Dataset
- 用模块化流量生成器模拟真实场景中的良性与恶意流量
- 构建含82维特征、21类细粒度标签的1200万样本数据集
- 重点覆盖常被忽略的网页攻击类型,适合安全研究与模型训练
基于异常的网络入侵检测系统(NIDS)需要标注准确、具有代表性且多样化的数据集以实现有效评估与开发。然而,现有广泛使用的数据集普遍存在标签粒度不足、样本量小的问题,易导致过拟合并难以在测试中发现。同时,网络安全领域发展迅速,新攻击手段不断出现,亟需持续更新数据集。为此,我们设计了一种模块化流量生成器,可模拟多种协议、通过随机化技术引入多样性,并在真实场景中同步生成对应良性与恶意流量。利用该生成器,我们构建了包含超过1200万样本的数据集,具有82个流级特征和21个细粒度标签,涵盖多种常被低估的网页攻击类型。
原文摘要 · Abstract (English)
Anomaly-based Network Intrusion Detection Systems (NIDS) require correctly labelled, representative and diverse datasets for an accurate evaluation and development. However, several widely used datasets do not include labels which are fine-grained enough and, together with small sample sizes, can lead to overfitting issues that also remain undetected when using test data. Additionally, the cybersecurity sector is evolving fast, and new attack mechanisms require the continuous creation of up-to-date datasets. To address these limitations, we developed a modular traffic generator that can simulate a wide variety of benign and malicious traffic. It incorporates multiple protocols, variability through randomization techniques and can produce attacks along corresponding benign traffic, as it occurs in real-world scenarios. Using the traffic generator, we create a dataset capturing over 12 million samples with 82 flow-level features and 21 fine-grained labels. Additionally, we include several web attack types which are often underrepresented in other datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。