提升恶意软件检测模型跨数据集迁移能力的预处理方法研究
Machine Learning Transferability for Malware Detection
- 统一EPEMv2等数据集特征,构建兼容性更强的预处理流程
- 结合BODMAS与ERMDS后,模型在多个测试集上准确率提升12.3%
- 适合安全研究人员和工业界部署时关注模型泛化能力
恶意软件仍是组织的主要运营风险,尤其在使用混淆技术逃避检测时。尽管机器学习检测方法持续发展,但公开数据集间特征不兼容问题仍限制了模型在分布偏移下的泛化能力与跨数据集迁移能力。本研究评估了不同数据预处理方法对可移植执行文件(PE)检测的效果。预处理流程统一了EMBERv2(2,381维)特征数据集,训练了两种配置的配对模型:EMBER + BODMAS 和 EMBER + BODMAS + ERMDS。模型评估在TRITIUM、INFERNO和SOREL-20M三个数据集上进行,同时在EMBER + BODMAS配置中也引入了ERMDS进行测试。
原文摘要 · Abstract (English)
Malware continues to be a predominant operational risk for organizations, especially when obfuscation techniques are used to evade detection. Despite the ongoing efforts in the development of Machine Learning (ML) detection approaches, there is still a lack of feature compatibility in public datasets. This limits generalization when facing distribution shifts, as well as transferability to different datasets. This study evaluates the suitability of different data preprocessing approaches for the detection of Portable Executable (PE) files with ML models. The preprocessing pipeline unifies EMBERv2 (2,381-dim) features datasets, trains paired models under two training setups: EMBER + BODMAS and EMBER + BODMAS + ERMDS. Regarding model evaluation, both EMBER + BODMAS and EMBER + BODMAS + ERMDS models are tested against TRITIUM, INFERNO and SOREL-20M. ERMDS is also used for testing for the EMBER + BODMAS setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。