arXiv:2511.08660cs.CRcs.AI2025-11

用GeNIS数据集验证了AI在网络安全检测中的可靠性。

Binary and Multiclass Cyberattack Classification on GeNIS Dataset

  • 结合五种特征选择法降维,提升检测效率
  • 机器学习模型在准确率和泛化上略胜深度学习
  • 适合做网络入侵检测的基准测试参考

人工智能在网络安全检测系统(NIDS)中的应用有望应对日益复杂的网络攻击。然而,由于机器学习和深度学习模型高度依赖训练数据质量,缺乏多样且更新及时的数据集限制了其对未见过的网络流量中恶意行为的泛化能力。本研究通过实验验证了GeNIS数据集在基于AI的NIDS中的可靠性,可作为未来基准的参考。采用信息增益、卡方检验、递归特征消除、平均绝对偏差和离散比率五种特征选择方法,筛选出GeNIS数据集中最相关特征并降低维度,以实现更高效的检测。训练了三种决策树集成模型和两种深度神经网络,分别用于二分类与多分类任务。所有模型均达到高准确率与高F1分数,其中机器学习集成模型在泛化性能上略优,同时计算效率更高。总体结果表明,GeNIS数据集支持基于时间与数量行为特征的智能入侵检测与攻击分类。

原文摘要 · Abstract (English)

The integration of Artificial Intelligence (AI) in Network Intrusion Detection Systems (NIDS) is a promising approach to tackle the increasing sophistication of cyberattacks. However, since Machine Learning (ML) and Deep Learning (DL) models rely heavily on the quality of their training data, the lack of diverse and up-to-date datasets hinders their generalization capability to detect malicious activity in previously unseen network traffic. This study presents an experimental validation of the reliability of the GeNIS dataset for AI-based NIDS, to serve as a baseline for future benchmarks. Five feature selection methods, Information Gain, Chi-Squared Test, Recursive Feature Elimination, Mean Absolute Deviation, and Dispersion Ratio, were combined to identify the most relevant features of GeNIS and reduce its dimensionality, enabling a more computationally efficient detection. Three decision tree ensembles and two deep neural networks were trained for both binary and multiclass classification tasks. All models reached high accuracy and F1-scores, and the ML ensembles achieved slightly better generalization while remaining more efficient than DL models. Overall, the obtained results indicate that the GeNIS dataset supports intelligent intrusion detection and cyberattack classification with time-based and quantity-based behavioral features.

网络安全特征选择入侵检测机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。