用决策树规则检测恶意软件演化中的概念漂移,提升分类稳定性。
Detecting Concept Drift in Evolving Malware Families Using Rule-Based Classifier Representations

- 通过提取决策树规则并比较特征重要性、预测一致性和覆盖度来量化漂移。
- 固定两个月窗口配合特征相关性分析,对所有家族对均呈现正向漂移-准确率关联。
- 适用于需要长期监控恶意软件演化的安全系统,尤其适合对抗持续变异的家族。
本文提出一种基于结构的恶意软件分类中概念漂移检测方法,利用决策树规则集。在EMBER2024数据集上,于多个时间窗口训练分类器,并通过特征重要性、预测一致性、激活稳定性与覆盖率等指标比较提取的规则表示,以量化漂移。这些指标与准确率下降和数据分布变化呈互补关系。在六类恶意软件中,采用固定间隔与基于聚类的时间窗口,在家族对良性及家族对家族场景下进行评估,并与RIPPER和Transcendent基线对比。结果表明,固定两个月窗口结合特征级皮尔逊相关性为最可靠配置,唯一实现所有家族对的正向漂移-准确率相关性。各方法表现互补,无单一方案在所有情况下占优。
原文摘要 · Abstract (English)
This work proposes a structural approach to concept drift detection in malware classification using decision tree rulesets. Classifiers are trained across temporal windows on the EMBER2024 dataset, and drift is quantified by comparing extracted rule representations using feature importance, prediction agreement, activation stability, and coverage metrics. These metrics are correlated with both accuracy degradation and data distribution shift as complementary drift indicators. The approach is evaluated across six malware families using fixed-interval and clustering-based windowing in family-vs-benign and family-vs-family settings, and compared against RIPPER and Transcendent baselines. Results show that fixed two-month windowing with feature-level Pearson correlation is the most reliable configuration, being the only one where all family pairs produce positive drift-accuracy correlations. The methods are complementary - no single approach dominates across all pairs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。