综述5大数据集在机器学习入侵检测中的应用与挑战
A Review of Various Datasets for Machine Learning Algorithm-Based Intrusion Detection System: Advances and Challenges
- 系统梳理KDDCUP'99等5个主流数据集在入侵检测中的使用方法
- 对比10种算法在不同数据集上的检测准确率与适用场景
- 适合安全研究者参考数据集选型与模型评估策略
入侵检测系统(IDS)通过识别、告警和响应来防范网络攻击,保护信息安全。随着全球对技术依赖加深,保障系统与网络的安全成为当务之急。近年来,研究者利用多种机器学习与深度学习模型提升检测能力。本文回顾过去十年基于机器学习的入侵检测研究,重点分析了KDDCUP'99、NSL-KDD、UNSW-NB15、CICIDS-2017和CSE-CIC-IDS2018等数据集的应用情况。文章详细评述了支持向量机(SVM)、K近邻(KNN)、决策树(DT)、逻辑回归(LR)、朴素贝叶斯(NB)、随机森林(RF)、XGBoost、Adaboost及人工神经网络(ANN)等算法的表现。通过表格形式对比各数据集、分类器、攻击类型、评估指标及研究结论,为未来研究提供全面参考。
原文摘要 · Abstract (English)
IDS aims to protect computer networks from security threats by detecting, notifying, and taking appropriate action to prevent illegal access and protect confidential information. As the globe becomes increasingly dependent on technology and automated processes, ensuring secured systems, applications, and networks has become one of the most significant problems of this era. The global web and digital technology have significantly accelerated the evolution of the modern world, necessitating the use of telecommunications and data transfer platforms. Researchers are enhancing the effectiveness of IDS by incorporating popular datasets into machine learning algorithms. IDS, equipped with machine learning classifiers, enhances security attack detection accuracy by identifying normal or abnormal network traffic. This paper explores the methods of capturing and reviewing intrusion detection systems (IDS) and evaluates the challenges existing datasets face. A deluge of research on machine learning (ML) and deep learning (DL) architecture-based intrusion detection techniques has been conducted in the past ten years on various cybersecurity datasets, including KDDCUP'99, NSL-KDD, UNSW-NB15, CICIDS-2017, and CSE-CIC-IDS2018. We conducted a literature review and presented an in-depth analysis of various intrusion detection methods that use SVM, KNN, DT, LR, NB, RF, XGBOOST, Adaboost, and ANN. We provide an overview of each technique, explaining the role of the classifiers and algorithms used. A detailed tabular analysis highlights the datasets used, classifiers employed, attacks detected, evaluation metrics, and conclusions drawn. This article offers a thorough review for future IDS research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。