用无标签数据提升网络安全中的威胁检测能力
Applications of Positive Unlabeled (PU) and Negative Unlabeled (NU) Learning in Cybersecurity
- 利用正样本和无标签数据训练模型,解决标注数据稀缺问题
- 在入侵检测等场景中显著提升对未知威胁的识别率
- 适合安全团队资源有限时快速构建检测系统
本文探讨了正样本-无标签(PU)学习与负样本-无标签(NU)学习在网络安全领域的应用潜力。尽管这些半监督学习方法已在医疗、营销等领域成功应用,但在网络安全中仍鲜有研究。论文指出,入侵检测、漏洞管理、恶意软件检测和威胁情报等关键领域,尤其适合使用此类方法应对标注数据稀少、类别不平衡的问题。文中为各子领域提供详细问题建模与数学分析,揭示其在实时系统部署、对抗动态威胁及处理样本不均衡方面的挑战,并提出未来研究方向,以推动该技术在新型网络威胁发现与防御中的深度整合。
原文摘要 · Abstract (English)
This paper explores the relatively underexplored application of Positive Unlabeled (PU) Learning and Negative Unlabeled (NU) Learning in the cybersecurity domain. While these semi-supervised learning methods have been applied successfully in fields like medicine and marketing, their potential in cybersecurity remains largely untapped. The paper identifies key areas of cybersecurity--such as intrusion detection, vulnerability management, malware detection, and threat intelligence--where PU/NU learning can offer significant improvements, particularly in scenarios with imbalanced or limited labeled data. We provide a detailed problem formulation for each subfield, supported by mathematical reasoning, and highlight the specific challenges and research gaps in scaling these methods to real-time systems, addressing class imbalance, and adapting to evolving threats. Finally, we propose future directions to advance the integration of PU/NU learning in cybersecurity, offering solutions that can better detect, manage, and mitigate emerging cyber threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。