arXiv:2512.23987cs.CRcs.AI2025-12

用元学习和分块特征选择,提升恶意软件检测的适应性和效率

MeLeMaD: Adaptive Malware Detection via Chunk-wise Feature Selection and Meta-Learning

  • 结合元学习与梯度提升的分块特征选择,高效处理高维恶意软件数据
  • 在三个数据集上准确率最高达99.97%,显著优于现有方法
  • 适合需要快速适应新型恶意软件的网络安全系统研发者

面对网络安全中恶意软件检测面临的巨大挑战,亟需具备鲁棒性与适应性的解决方案。本文提出基于模型无关元学习(MAML)的恶意软件检测框架MeLeMaD,引入一种针对大规模高维恶意软件数据集设计的新型特征选择技术——基于梯度提升的分块特征选择(CFSGB),显著提升检测效率。通过CIC-AndMal2020、BODMAS两个基准数据集及自建数据集EMBOD进行严格验证,结果表明,在准确率、精确率、召回率、F1分数、马修相关系数(MCC)和AUC等关键指标上均表现优异:在CIC-AndMal2020上准确率达98.04%,BODMAS上达99.97%,EMBOD上达97.85%。实验充分证明了MeLeMaD在应对鲁棒性、适应性及大规模高维数据挑战方面的潜力,为更高效可靠的网络安全防护提供新路径。

原文摘要 · Abstract (English)

Confronting the substantial challenges of malware detection in cybersecurity necessitates solutions that are both robust and adaptable to the ever-evolving threat environment. The paper introduces Meta Learning Malware Detection (MeLeMaD), a novel framework leveraging the adaptability and generalization capabilities of Model-Agnostic Meta-Learning (MAML) for malware detection. MeLeMaD incorporates a novel feature selection technique, Chunk-wise Feature Selection based on Gradient Boosting (CFSGB), tailored for handling large-scale, high-dimensional malware datasets, significantly enhancing the detection efficiency. Two benchmark malware datasets (CIC-AndMal2020 and BODMAS) and a custom dataset (EMBOD) were used for rigorously validating the MeLeMaD, achieving a remarkable performance in terms of key evaluation measures, including accuracy, precision, recall, F1-score, MCC, and AUC. With accuracies of 98.04\% on CIC-AndMal2020 and 99.97\% on BODMAS, MeLeMaD outperforms the state-of-the-art approaches. The custom dataset, EMBOD, also achieves a commendable accuracy of 97.85\%. The results underscore the MeLeMaD's potential to address the challenges of robustness, adaptability, and large-scale, high-dimensional datasets in malware detection, paving the way for more effective and efficient cybersecurity solutions.

恶意软件检测元学习特征选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。