arXiv:2511.20480cs.LGcs.AI2025-11被引 4

用主动学习提升自编码器检测隐蔽网络攻击的能力

Ranking-Enhanced Anomaly Detection Using Active Learning-Assisted Attention Adversarial Dual AutoEncoders

  • 结合注意力对抗双自编码器与主动学习迭代优化
  • 在0.004%极低异常占比下检测率显著提升
  • 适合数据稀缺的网络安全场景,降低人工标注成本

高级持续性威胁(APTs)因其隐秘性和长期性给网络安全带来重大挑战。现代监督学习方法依赖大量标注数据,但在真实环境中往往难以获取。本文提出一种基于自编码器的无监督异常检测方法,通过主动学习机制有选择地向人工标注者查询不确定样本标签,以最小化标注成本并持续提升检测精度。该框架采用注意力对抗双自编码器结构,并在多系统(Android、Linux、BSD、Windows)和双攻击场景的真实溯源数据集上进行评估,其中APT类攻击仅占0.004%。实验结果表明,主动学习循环显著提升了检测性能,优于现有主流方法。

原文摘要 · Abstract (English)

Advanced Persistent Threats (APTs) pose a significant challenge in cybersecurity due to their stealthy and long-term nature. Modern supervised learning methods require extensive labeled data, which is often scarce in real-world cybersecurity environments. In this paper, we propose an innovative approach that leverages AutoEncoders for unsupervised anomaly detection, augmented by active learning to iteratively improve the detection of APT anomalies. By selectively querying an oracle for labels on uncertain or ambiguous samples, we minimize labeling costs while improving detection rates, enabling the model to improve its detection accuracy with minimal data while reducing the need for extensive manual labeling. We provide a detailed formulation of the proposed Attention Adversarial Dual AutoEncoder-based anomaly detection framework and show how the active learning loop iteratively enhances the model. The framework is evaluated on real-world imbalanced provenance trace databases produced by the DARPA Transparent Computing program, where APT-like attacks constitute as little as 0.004\% of the data. The datasets span multiple operating systems, including Android, Linux, BSD, and Windows, and cover two attack scenarios. The results have shown significant improvements in detection rates during active learning and better performance compared to other existing approaches.

异常检测主动学习网络安全自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。