构建可绕过杀毒软件的恶意软件数据集,用于测试检测系统鲁棒性。
Building an Adversarial Malware Dataset by Family and Type: Generation, Evasion, and Poisoning Evaluation

- 用生成器创建4.4万条带家族标签的对抗样本,98.35%可绕过分类器。
- 仅0.5%污染样本使重训练模型绕过率从26.1%升至92.8%。
- 适合研究对抗攻击、数据投毒和机器学习杀毒系统安全的人使用。
我们基于公开的RawMal-TF真实恶意软件二进制文件,利用一系列对抗性恶意软件生成器,构建了两组对抗性PE文件:44,347个带家族标签的样本和33,596个带类型标签的样本,分别在EMBER分类器上实现了98.35%和92.20%的逃逸率。每个对抗样本均附有详细元数据,包括EMBER评分和VirusTotal分类结果。通过一系列训练实验,我们进一步验证了恶意软件分类流水线对数据投毒攻击的脆弱性:在家族标签数据集中,仅注入0.5%完全错误标注的对抗样本,即可使重训练分类器的逃逸率从26.1%提升至92.8%。该数据集已公开发布,以促进未来对抗性恶意软件、投毒攻击及机器学习驱动的恶意软件检测系统鲁棒性研究。
原文摘要 · Abstract (English)
We present a dataset of adversarial malware samples derived from the public RawMal-TF collection of real-world malware binaries. Using a suite of adversarial malware generators, we construct two sets of adversarial PE files: 44,347 family-labelled samples and 33,596 type-labelled samples, achieving evasion rates of 98.35 % and 92.20 % against the EMBER classifier, respectively. Each adversarial binary is accompanied by detailed metadata, including EMBER scores and VirusTotal classifications. We further demonstrate the susceptibility of malware classification pipelines to data poisoning attacks through a series of training experiments. Injecting fully mislabelled adversarial samples representing only 0.5 % of the training data in the family-labelled dataset increases the evasion rate against the re-trained classifier from 26.1 % to 92.8 %. The dataset is publicly released to facilitate future research on adversarial malware, poisoning attacks, and the robustness of machine-learning-based malware detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。