用大模型自动把汽车漏洞文本转成结构化威胁数据,提升安全响应效率。
Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities

- 用开源大模型将汽车漏洞描述转化为STIX格式,实现威胁信息结构化
- 单模型最高达0.94的SDO识别准确率,CWE映射接近完美(0.99)
- 发现车联网漏洞中常见的攻击模式,适合安全团队做自动化威胁分析
联网与自动驾驶汽车依赖传感器、电子控制单元、车载信息娱乐系统和远程信息处理单元等软硬件组件,其漏洞可能危及资产、用户和车辆运行。这些漏洞通常以纯文本形式记录在通用漏洞披露(CVE)数据库中,但安全人员需要结构化信息来识别受影响资产、弱点类型和攻击行为,以有效缓解风险。为此,本文评估开源大语言模型(LLM)生成结构化威胁信息表达(STIX)的能力,用于车联网相关CVE。我们构建了名为CAV-STIXGen的数据集,将车联网漏洞描述映射到STIX域对象(SDO)、关系对象(SRO)、通用弱点枚举(CWE)及MITRE ATT&CK技术。基于该数据集,评估了11个参数量从4B到120B的开源LLM,涵盖多种提示策略与温度设置。单模型配置下,SDO、SRO和CWE映射的F1分数分别为0.94、0.63和0.99;完整MITRE ATT&CK映射仍具挑战性。多智能体设置中,Gemma-4-31B与Codestral-22B分别达到SDO F1为0.91、SRO F1为0.43。最后,分析CWE与MITRE ATT&CK共现关系,揭示车联网领域中反复出现的威胁模式,表明AI辅助的漏洞到STIX转换可实现威胁情报自动化并优先防御交通安全隐患。
原文摘要 · Abstract (English)
Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Units, in-vehicle infotainment systems, and telematics units, where vulnerabilities can compromise assets, users, and vehicle operations. These vulnerabilities are commonly documented as plain text in the Common Vulnerabilities and Exposures (CVE) database; however, security practitioners require structured information about affected assets, types of weaknesses, and attack behaviors to effectively mitigate the risks from these vulnerabilities. To this end, we evaluate open-weight Large Language Models (LLMs) for generating Structured Threat Information Expression (STIX), a well-known structured format for representing threat information, for CAV-related CVEs. We construct a dataset called CAV-STIXGen that maps CAV vulnerability descriptions to STIX domain objects (SDO), STIX relationship objects (SRO), Common Weakness Enumeration (CWE), and MITRE ATT&CK techniques mappings. Using this dataset, we evaluated 11 open-weight LLMs (4B to 120B parameters) across various prompting strategies and temperatures. Single-model configurations achieve F1 scores of 0.94 for SDO, 0.63 for SRO, and 0.99 for CWE mapping, while complete MITRE ATT&CK mapping remains challenging. In a multi-agent setup, Gemma-4-31B and Codestral-22B achieve F1 scores of 0.91 for SDOs and 0.43 for SROs, respectively. Lastly, we analyze CWE and MITRE ATT&CK co-occurrences to identify recurring threat patterns in the CAV domain, demonstrating how AI-assisted vulnerability-to-STIX translation can automate threat intelligence and prioritize defense in transportation security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。