arXiv:2412.06203cs.CRcs.LG2024-12

用无标签数据提升网络安全中的威胁检测能力

Applications of Positive Unlabeled (PU) and Negative Unlabeled (NU) Learning in Cybersecurity

  • 利用正样本和无标签数据训练模型,解决标注数据稀缺问题
  • 在入侵检测等场景中显著提升对未知威胁的识别率
  • 适合安全团队资源有限时快速构建检测系统

本文探讨了正样本-无标签(PU)学习与负样本-无标签(NU)学习在网络安全领域的应用潜力。尽管这些半监督学习方法已在医疗、营销等领域成功应用,但在网络安全中仍鲜有研究。论文指出,入侵检测、漏洞管理、恶意软件检测和威胁情报等关键领域,尤其适合使用此类方法应对标注数据稀少、类别不平衡的问题。文中为各子领域提供详细问题建模与数学分析,揭示其在实时系统部署、对抗动态威胁及处理样本不均衡方面的挑战,并提出未来研究方向,以推动该技术在新型网络威胁发现与防御中的深度整合。

原文摘要 · Abstract (English)

This paper explores the relatively underexplored application of Positive Unlabeled (PU) Learning and Negative Unlabeled (NU) Learning in the cybersecurity domain. While these semi-supervised learning methods have been applied successfully in fields like medicine and marketing, their potential in cybersecurity remains largely untapped. The paper identifies key areas of cybersecurity--such as intrusion detection, vulnerability management, malware detection, and threat intelligence--where PU/NU learning can offer significant improvements, particularly in scenarios with imbalanced or limited labeled data. We provide a detailed problem formulation for each subfield, supported by mathematical reasoning, and highlight the specific challenges and research gaps in scaling these methods to real-time systems, addressing class imbalance, and adapting to evolving threats. Finally, we propose future directions to advance the integration of PU/NU learning in cybersecurity, offering solutions that can better detect, manage, and mitigate emerging cyber threats.

网络安全半监督学习威胁检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。