用语言模型识别漏洞间的父子关系,构建层级攻击模型。
Towards the generation of hierarchical attack models from cybersecurity vulnerabilities using language models
- 基于预训练语言模型与孪生网络判断漏洞文本的上下级关系。
- 提出两种采样策略和共识机制,降低误报率并提升预测稳定性。
- 在三个真实数据集上验证方法,适用于安全分析与威胁建模场景。
本文研究使用预训练语言模型与孪生网络,从基于文本的网络安全漏洞数据中识别兄弟关系。目标是基于对系统潜在或已观察漏洞的文本描述,构建层级攻击模型。由于数据特性及问题环境的不确定性,需采用实用的软计算方法。因此,本工作重点探讨预测链接在构建模型中的可靠性,提出应对数据复杂性与预测不稳定性等概念与实践挑战的解决方案。主要贡献包括:利用预训练语言模型构建神经网络以预测漏洞间的兄弟关系,并阐述如何将该能力应用于生成层级攻击模型;提出两种数据采样机制以缓解数据复杂性,以及一种共识机制以减少误报;上述方法在三个网络安全数据集上通过实证结果进行对比分析,评估其有效性。
原文摘要 · Abstract (English)
This paper investigates the use of a pre-trained language model and siamese network to discern sibling relationships between text-based cybersecurity vulnerability data. The ultimate purpose of the approach presented in this paper is towards the construction of hierarchical attack models based on a set of text descriptions characterising potential/observed vulnerabilities in a given system. Due to the nature of the data, and the uncertainty sensitive environment in which the problem is presented, a practically oriented soft computing approach is necessary. Therefore, a key focus of this work is to investigate practical questions surrounding the reliability of predicted links towards the construction of such models, to which end conceptual and practical challenges and solutions associated with the proposed approach are outlined, such as dataset complexity and stability of predictions. Accordingly, the contributions of this paper focus on producing neural networks using a pre-trained language model for predicting sibling relationships between cybersecurity vulnerabilities, then outlining how to apply this capability towards the generation of hierarchical attack models. In addition, two data sampling mechanisms for tackling data complexity, and a consensus mechanism for reducing the amount of false positive predictions are outlined. Each of these approaches is compared and contrasted using empirical results from three sets of cybersecurity data to determine their effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。