构建攻击序列理解基准,评估大模型在网络安全中的推理能力。
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
- 设计多维度攻击行为评测框架,覆盖战术、技术与流程层面。
- 测试7个大模型在3个任务中表现,揭示其在序列理解上的短板。
- 适合安全研究者和大模型应用开发者参考使用。
网络威胁情报(CTI)报告记录了对网络威胁的观察,将攻击者行动与意图的证据整合为可操作的知识,用于检测、响应和防御规划。然而,CTI报告内容非结构化且冗长,给安全人员手动提取分析带来挑战。尽管大语言模型(LLMs)在实体抽取和知识图谱构建等任务中展现出潜力,但其对攻击行为序列的理解与推理能力仍待深入探索。为此,我们提出AttackSeqBench,一个旨在系统评估LLMs在战术、技术与程序维度上推理能力的基准,具备可扩展性、推理可扩展性及领域认知可拓展性。我们在三个基准设置和三个任务中,对7个LLMs、5个LRMs及4种后训练策略进行了评测,识别出其在特定领域的优劣。研究结果深化了对大模型驱动的CTI报告理解的认识,并推动其在网络安全运营中的应用。基准构建与评估代码及数据集已开源:https://github.com/hulkima/AttackSeqBench。
原文摘要 · Abstract (English)
Cyber Threat Intelligence (CTI) reports document observations of cyber threats, synthesizing evidence about adversaries' actions and intent into actionable knowledge that informs detection, response, and defense planning. However, the unstructured and verbose nature of CTI reports poses significant challenges for security practitioners to manually extract and analyze such sequences. Although large language models (LLMs) exhibit promise in cybersecurity tasks such as entity extraction and knowledge graph construction, their understanding and reasoning capabilities towards behavioral sequences remains underexplored. To address this, we introduce AttackSeqBench, a benchmark designed to systematically evaluate LLMs' reasoning abilities across the tactical, technical, and procedural dimensions of adversarial behaviors, while satisfying Extensibility, Reasoning Scalability, and Domain-dpecific Epistemic Expandability. We further benchmark 7 LLMs, 5 LRMs and 4 post-training strategies across 3 benchmark settings and 3 benchmark tasks within our AttackSeqBench to identify their advantages and limitations in such specific domain. Our findings contribute to a deeper understanding of LLM-driven CTI report understanding and foster its application in cybersecurity operations. Our code of benchmark construction and evaluation and the corresponding dataset are available at: https://github.com/hulkima/AttackSeqBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。