用真实攻击场景数据训练模型,提升对IPv6隐蔽通信的检测准确率。
AI/ML Based Detection and Categorization of Covert Communication in IPv6 Network
- 构建贴近真实攻击的IPv6隐蔽通信数据集,保持网络行为不变。
- 多模型对比验证,检测准确率超90%。
- 引入生成式AI辅助检测与模型优化,探索新方法。
IPv6扩展头的灵活性和复杂性使攻击者可创建隐蔽通道或绕过安全机制,导致数据泄露或系统被攻陷。机器学习已成为应对隐蔽通信威胁的主要技术手段,但检测复杂度高、攻击手法不断演进及数据稀缺,使模型构建面临挑战。以往研究虽在检测中表现良好,但攻击场景假设过于简化,无法反映现代隐蔽技术的复杂性,反而让模型容易识别。为此,本研究分析了IPv6报文结构与网络流量行为,采用加密算法实现隐蔽通信注入,且不改变网络包行为,使攻击更贴近真实场景。在此基础上,结合随机森林、梯度提升等传统决策树与卷积神经网络(CNN)、长短期记忆网络(LSTM)等复杂神经网络,训练检测模型,实现检测准确率超过90%。研究还详细阐述了数据增强方法及各模型性能对比,揭示了机器学习在应对复杂隐蔽通信中的适应性与鲁棒性。此外,引入基于生成式AI的脚本优化框架,通过提示工程初步探索生成智能体在隐蔽通信检测与模型增强中的应用潜力。
原文摘要 · Abstract (English)
The flexibility and complexity of IPv6 extension headers allow attackers to create covert channels or bypass security mechanisms, leading to potential data breaches or system compromises. The mature development of machine learning has become the primary detection technology option used to mitigate covert communication threats. However, the complexity of detecting covert communication, evolving injection techniques, and scarcity of data make building machine-learning models challenging. In previous related research, machine learning has shown good performance in detecting covert communications, but oversimplified attack scenario assumptions cannot represent the complexity of modern covert technologies and make it easier for machine learning models to detect covert communications. To bridge this gap, in this study, we analyzed the packet structure and network traffic behavior of IPv6, used encryption algorithms, and performed covert communication injection without changing network packet behavior to get closer to real attack scenarios. In addition to analyzing and injecting methods for covert communications, this study also uses comprehensive machine learning techniques to train the model proposed in this study to detect threats, including traditional decision trees such as random forests and gradient boosting, as well as complex neural network architectures such as CNNs and LSTMs, to achieve detection accuracy of over 90\%. This study details the methods used for dataset augmentation and the comparative performance of the applied models, reinforcing insights into the adaptability and resilience of the machine learning application in IPv6 covert communication. We further introduce a Generative AI-driven script refinement framework, leveraging prompt engineering as a preliminary exploration of how generative agents can assist in covert communication detection and model enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。