首次在全量数据集上评估恶意软件聚类,发现主流方法效果优于预期。
Clustering Malware at Scale: A First Full-Benchmark Study
- 在Bodmas和Ember全量数据集上系统评估聚类性能
- 引入良性样本后聚类质量未明显下降,且不同数据集表现差异大
- K-Means与BIRCH表现最佳,非主流的DBSCAN/HAC反而落后
近年来恶意软件攻击频发,专家需对样本分类以判断其可信性或恶意性。恶意软件聚类是识别样本群体的重要手段。然而,现有研究多忽略良性样本的影响,且普遍使用小规模数据集(仅包含少数家族),未能充分利用大型公开基准数据集。本文首次在完整规模的Bodmas与Ember两个公开数据集上开展恶意软件聚类研究,并扩展任务以纳入良性样本。结果表明,加入良性样本并未显著降低聚类质量;同时,Ember与Bodmas及某工业私有数据集上的聚类效果存在差异。出人意料的是,当前表现最佳的算法是K-Means与BIRCH,而传统认为性能优越的DBSCAN与层次聚类(HAC)则表现较弱。
原文摘要 · Abstract (English)
Recent years have shown that malware attacks still happen with high frequency. Malware experts seek to categorize and classify incoming samples to confirm their trustworthiness or prove their maliciousness. One of the ways in which groups of malware samples can be identified is through malware clustering. Despite the efforts of the community, malware clustering which incorporates benign samples has been under-explored. Moreover, despite the availability of larger public benchmark malware datasets, malware clustering studies have avoided fully utilizing these datasets in their experiments, often resorting to small datasets with only a few families. Additionally, the current state-of-the-art solutions for malware clustering remain unclear. In our study, we evaluate malware clustering quality and establish the state-of-the-art on Bodmas and Ember - two large public benchmark malware datasets. Ours is the first study of malware clustering performed on whole malware benchmark datasets. Additionally, we extend the malware clustering task by incorporating benign samples. Our results indicate that incorporating benign samples does not significantly degrade clustering quality. We find that there are differences in the quality of the created clusters between Ember and Bodmas, as well as a private industry dataset. Contrary to popular opinion, our top clustering performers are K-Means and BIRCH, with DBSCAN and HAC falling behind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。