用机器学习自动从威胁情报中识别漏洞被利用的事件。
Harnessing TI Feeds for Exploitation Detection
- 用Doc2Vec和BERT建模威胁术语,训练分类器识别漏洞利用。
- 在191个威胁情报源上准确识别出漏洞利用事件。
- 仅靠历史数据就能检测未参与训练的来源,适合安全分析场景。
许多组织依赖威胁情报(TI)源来评估安全威胁的风险。由于数据量大且格式松散不一,手动分析不同结构松散的TI源中的威胁信息成本过高。因此,亟需开发自动化方法,从TI源中筛选并提取可操作的信息。为此,我们提出一个机器学习流程,可自动从TI源中检测漏洞利用事件。首先,使用先进的嵌入技术(Doc2Vec和BERT)对松散结构的TI源中的威胁词汇进行建模,然后基于此训练监督式机器学习分类器以检测安全漏洞的利用行为。我们利用该方法在191个不同的TI源中识别漏洞利用事件。纵向评估表明,该方法仅使用历史数据进行训练,即可准确识别来自未参与训练的TI源的利用事件。本方法适用于多种下游任务,如数据驱动的漏洞风险评估。
原文摘要 · Abstract (English)
Many organizations rely on Threat Intelligence (TI) feeds to assess the risk associated with security threats. Due to the volume and heterogeneity of data, it is prohibitive to manually analyze the threat information available in different loosely structured TI feeds. Thus, there is a need to develop automated methods to vet and extract actionable information from TI feeds. To this end, we present a machine learning pipeline to automatically detect vulnerability exploitation from TI feeds. We first model threat vocabulary in loosely structured TI feeds using state-of-the-art embedding techniques (Doc2Vec and BERT) and then use it to train a supervised machine learning classifier to detect exploitation of security vulnerabilities. We use our approach to identify exploitation events in 191 different TI feeds. Our longitudinal evaluation shows that it is able to accurately identify exploitation events from TI feeds only using past data for training and even on TI feeds withheld from training. Our proposed approach is useful for a variety of downstream tasks such as data-driven vulnerability risk assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。