提出无模型依赖的更新框架,90%降标注成本仍保检测效果。
Label-efficient Training Updates for Malware Detection over Time
- 构建通用框架,独立评估主动与半监督学习
- 结合方法可降90%标注成本,性能媲美全量重训
- 引入特征级漂移分析,揭示性能变化原因
基于机器学习的恶意软件检测系统日益重要,但传统算法难以应对真实环境中合法与恶意软件的持续演化。分布漂移导致静态训练模型随时间退化,需频繁更新。然而,重新训练成本高昂,因新数据标注需安全专家人工分析。为降低标注成本并应对漂移,现有研究采用主动学习(AL)和半监督学习(SSL),但存在两大局限:(i)方法与特定检测器架构绑定,仅限特定恶意软件领域,难以统一比较;(ii)缺乏一致的分布漂移分析方法,而恶意软件领域对时间变化极为敏感。本文提出一种模型无关框架,系统评估安卓与Windows平台下多种AL与SSL技术的独立及组合效果。结果表明,组合策略可在两个领域将人工标注成本降低高达90%,同时保持与全量标注重训相当的检测性能。此外,提出特征级漂移分析方法,量化特征稳定性,发现其与检测性能显著相关。本研究揭示了AL与SSL在分布漂移下的行为机制,为长期有效检测器设计提供实用指导。
原文摘要 · Abstract (English)
Machine Learning (ML)-based detectors are becoming essential to counter the proliferation of malware. However, common ML algorithms are not designed to cope with the dynamic nature of real-world settings, where both legitimate and malicious software evolve. This distribution drift causes models trained under static assumptions to degrade over time unless they are continuously updated. Regularly retraining these models, however, is expensive, since labeling new acquired data requires costly manual analysis by security experts. To reduce labeling costs and address distribution drift in malware detection, prior work explored active learning (AL) and semi-supervised learning (SSL) techniques. Yet, existing studies (i) are tightly coupled to specific detector architectures and restricted to a specific malware domain, resulting in non-uniform comparisons; and (ii) lack a consistent methodology for analyzing the distribution drift, despite the critical sensitivity of the malware domain to temporal changes. In this work, we bridge this gap by proposing a model-agnostic framework that evaluates an extensive set of AL and SSL techniques, isolated and combined, for Android and Windows malware detection. We show that these techniques, when combined, can reduce manual annotation costs by up to 90% across both domains while achieving comparable detection performance to full-labeling retraining. We also introduce a methodology for feature-level drift analysis that measures feature stability over time, showing its correlation with the detector performance. Overall, our study provides a detailed understanding of how AL and SSL behave under distribution drift and how they can be successfully combined, offering practical insights for the design of effective detectors over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。