用大模型补全网络攻击知识图谱,提升安全风险评估能力。
Cyber Knowledge Completion Using Large Language Models
- 通过向量嵌入技术建立攻击模式与对抗手法的映射关系。
- 提出基于RAG的框架,实现跨威胁分类体系的结构化知识生成。
- 在小规模标注数据上验证效果优于传统分类模型,适合安全研究者使用。
物联网(IoT)融入网络物理系统(CPS)后,其网络攻击面扩大,新型复杂威胁频发,潜在漏洞易被利用。由于网络安全知识不完整且过时,对CPS的风险评估日益困难,亟需更精准的风险评估与缓解策略。以往方法依赖基于规则的自然语言处理工具来关联漏洞、弱点和攻击模式,而大型语言模型(LLM)的进展为增强攻击知识补全提供了新可能,其具备更强的推理、推断和摘要能力。本文采用嵌入模型封装攻击模式与对抗技术信息,利用向量嵌入生成二者间的映射关系;同时提出一种基于检索增强生成(RAG)的方法,借助预训练模型构建不同威胁分类体系间的结构化映射。此外,通过一个小规模人工标注数据集,将所提RAG方法与标准二分类基线模型进行对比。结果表明,该方法能有效完成网络攻击知识图谱补全,提供一个全面的解决方案。
原文摘要 · Abstract (English)
The integration of the Internet of Things (IoT) into Cyber-Physical Systems (CPSs) has expanded their cyber-attack surface, introducing new and sophisticated threats with potential to exploit emerging vulnerabilities. Assessing the risks of CPSs is increasingly difficult due to incomplete and outdated cybersecurity knowledge. This highlights the urgent need for better-informed risk assessments and mitigation strategies. While previous efforts have relied on rule-based natural language processing (NLP) tools to map vulnerabilities, weaknesses, and attack patterns, recent advancements in Large Language Models (LLMs) present a unique opportunity to enhance cyber-attack knowledge completion through improved reasoning, inference, and summarization capabilities. We apply embedding models to encapsulate information on attack patterns and adversarial techniques, generating mappings between them using vector embeddings. Additionally, we propose a Retrieval-Augmented Generation (RAG)-based approach that leverages pre-trained models to create structured mappings between different taxonomies of threat patterns. Further, we use a small hand-labeled dataset to compare the proposed RAG-based approach to a baseline standard binary classification model. Thus, the proposed approach provides a comprehensive framework to address the challenge of cyber-attack knowledge graph completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。