用元学习和分块特征选择,提升恶意软件检测的适应性和效率
MeLeMaD: Adaptive Malware Detection via Chunk-wise Feature Selection and Meta-Learning
- 结合元学习与梯度提升的分块特征选择,高效处理高维恶意软件数据
- 在三个数据集上准确率最高达99.97%,显著优于现有方法
- 适合需要快速适应新型恶意软件的网络安全系统研发者
面对网络安全中恶意软件检测面临的巨大挑战,亟需具备鲁棒性与适应性的解决方案。本文提出基于模型无关元学习(MAML)的恶意软件检测框架MeLeMaD,引入一种针对大规模高维恶意软件数据集设计的新型特征选择技术——基于梯度提升的分块特征选择(CFSGB),显著提升检测效率。通过CIC-AndMal2020、BODMAS两个基准数据集及自建数据集EMBOD进行严格验证,结果表明,在准确率、精确率、召回率、F1分数、马修相关系数(MCC)和AUC等关键指标上均表现优异:在CIC-AndMal2020上准确率达98.04%,BODMAS上达99.97%,EMBOD上达97.85%。实验充分证明了MeLeMaD在应对鲁棒性、适应性及大规模高维数据挑战方面的潜力,为更高效可靠的网络安全防护提供新路径。
原文摘要 · Abstract (English)
Confronting the substantial challenges of malware detection in cybersecurity necessitates solutions that are both robust and adaptable to the ever-evolving threat environment. The paper introduces Meta Learning Malware Detection (MeLeMaD), a novel framework leveraging the adaptability and generalization capabilities of Model-Agnostic Meta-Learning (MAML) for malware detection. MeLeMaD incorporates a novel feature selection technique, Chunk-wise Feature Selection based on Gradient Boosting (CFSGB), tailored for handling large-scale, high-dimensional malware datasets, significantly enhancing the detection efficiency. Two benchmark malware datasets (CIC-AndMal2020 and BODMAS) and a custom dataset (EMBOD) were used for rigorously validating the MeLeMaD, achieving a remarkable performance in terms of key evaluation measures, including accuracy, precision, recall, F1-score, MCC, and AUC. With accuracies of 98.04\% on CIC-AndMal2020 and 99.97\% on BODMAS, MeLeMaD outperforms the state-of-the-art approaches. The custom dataset, EMBOD, also achieves a commendable accuracy of 97.85\%. The results underscore the MeLeMaD's potential to address the challenges of robustness, adaptability, and large-scale, high-dimensional datasets in malware detection, paving the way for more effective and efficient cybersecurity solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。