用扩散模型生成少数类数据,提升网络攻击分类召回率。
Diffusion-Driven Synthetic Tabular Data Generation for Enhanced DoS/DDoS Attack Classification
- 基于TabDDPM迭代去噪生成高质量少数类样本。
- 增强后数据使ANN分类器对低频攻击的召回率接近完美。
- 适合安全、金融、医疗等领域处理数据不平衡问题。
类别不平衡指数据集中某些类别的样本数量远少于其他类别,导致模型性能偏差。本文针对使用表格去噪扩散概率模型(TabDDPM)进行数据增强时的网络入侵检测类别不平衡问题提出解决方案。通过迭代去噪过程,从CIC-IDS2017数据集合成高保真度的少数类样本,并将其合并至原始数据集。增强后的训练数据使ANN分类器在以往被低估的攻击类别上实现近乎完美的召回率。结果表明,扩散模型是安全领域表格式数据不平衡问题的有效解决方法,具有在欺诈检测和医学诊断中的应用潜力。
原文摘要 · Abstract (English)
Class imbalance refers to a situation where certain classes in a dataset have significantly fewer samples than oth- ers, leading to biased model performance. Class imbalance in network intrusion detection using Tabular Denoising Diffusion Probability Models (TabDDPM) for data augmentation is ad- dressed in this paper. Our approach synthesizes high-fidelity minority-class samples from the CIC-IDS2017 dataset through iterative denoising processes. For the minority classes that have smaller samples, synthetic samples were generated and merged with the original dataset. The augmented training data enables an ANN classifier to achieve near-perfect recall on previously underrepresented attack classes. These results establish diffusion models as an effective solution for tabular data imbalance in security domains, with potential applications in fraud detection and medical diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。