arXiv:2509.04999cs.CRcs.AI2025-09被引 2

用主动学习减少标注量,提升对极罕见网络攻击的检测能力

Adversarial Augmentation and Active Sampling for Robust Cyber Anomaly Detection

  • 结合自编码器与主动学习,仅需少量标注数据即可训练
  • 在0.004%稀有攻击数据上实现显著检测率提升
  • 适合数据稀缺但安全要求高的真实网络环境

高级持续性威胁(APTs)因其隐蔽性和长期性给网络安全带来巨大挑战。传统监督学习需要大量标注数据,但在实际场景中往往难以获取。本文提出一种新型方法,将自编码器用于异常检测,并引入主动学习机制,通过有选择地向标注者查询不确定样本的标签,降低标注成本的同时提升检测精度,使模型能在极少数据下有效学习,减少对人工标注的依赖。我们提出了基于注意力对抗双自编码器的异常检测框架,并展示了主动学习循环如何逐步提升模型性能。该框架在来自DARPA透明计算项目的实际、不平衡的进程追溯数据上进行评估,其中类似APT的攻击仅占0.004%。数据涵盖Android、Linux、BSD和Windows等多个操作系统,测试了两种攻击场景。结果表明,主动学习过程显著提升了检测率,优于现有方法。

原文摘要 · Abstract (English)

Advanced Persistent Threats (APTs) present a considerable challenge to cybersecurity due to their stealthy, long-duration nature. Traditional supervised learning methods typically require large amounts of labeled data, which is often scarce in real-world scenarios. This paper introduces a novel approach that combines AutoEncoders for anomaly detection with active learning to iteratively enhance APT detection. By selectively querying an oracle for labels on uncertain or ambiguous samples, our method reduces labeling costs while improving detection accuracy, enabling the model to effectively learn with minimal data and reduce reliance on extensive manual labeling. We present a comprehensive formulation of the Attention Adversarial Dual AutoEncoder-based anomaly detection framework and demonstrate how the active learning loop progressively enhances the model's performance. The framework is evaluated on real-world, imbalanced provenance trace data from the DARPA Transparent Computing program, where APT-like attacks account for just 0.004\% of the data. The datasets, which cover multiple operating systems including Android, Linux, BSD, and Windows, are tested in two attack scenarios. The results show substantial improvements in detection rates during active learning, outperforming existing methods.

异常检测主动学习网络安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。