arXiv:2606.18166cs.CRcs.LG2026-06被引 1

评测开源大模型在复杂威胁报告中的多标签攻击技术分类能力

Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports

  • 构建2076条人工标注的威胁报告句子数据集,覆盖114个ATT&CK技术
  • 最大模型仅达F1=0.22,参数量越大性能越好但整体仍不满足生产需求
  • 揭示当前大模型在真实威胁情报中表现不足,适合安全研究者参考

利用MITRE ATT&CK对网络威胁情报(CTI)进行分类对主动防御至关重要,但传统方法依赖大量人力。预大模型时代自动化虽提升效率,却难以处理非结构化报告中的复杂语言和多步骤攻击模式。大模型通过上下文推理改善了理解能力,但现有评估多基于简化单技术句,忽略真实报告复杂性,导致性能被高估。本文从83份复杂非结构化CTI报告中构建包含2,076条句子的标注数据集(1,281条含技术,795条不含),经六阶段标注流程映射至114个唯一ATT&CK技术,达到κ=0.68的标注者一致性。在此数据集上评估了7个参数量从8B到236B的开源LLM,采用不同提示策略与温度配置。最优模型微平均F1仅为0.22,确立了复杂非结构化报告上多标签ATT&CK分类的实证基线。参数规模与F1呈显著正相关,而提示策略和温度无显著提升效果。结果表明当前开源大模型尚不足以支撑生产级分类任务。本研究提供可复现的数据集、基准与结论,为未来CTI研究奠定基础。

原文摘要 · Abstract (English)

Classifying Cyber Threat Intelligence (CTI) using MITRE Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) is essential for proactive defense, but historically required extensive human effort. Pre-Large Language Model (LLM) automation sped up this process, but could not resolve the complex language and multi-step attack patterns found in unstructured CTI reports. LLMs addressed previous limitations by using contextual reasoning to understand unstructured text. However, current evaluations rely on simplified, single-technique sentences that ignore the complexity of real-world CTI reports, which often leads to inflated performance results. Consequently, the baseline performance of open-source LLMs on complex unstructured CTI reports remains unevaluated. To address this gap, we constructed a ground-truth dataset of 2,076 human-annotated sentences (1,281 technique-positive, 795 negative) from 83 complex unstructured CTI reports. These sentences were mapped to 114 unique ATT&CK techniques using a six-phase annotation process, achieving \k{appa} = 0.68 inter-annotator agreement. Using this dataset, we evaluated seven open-source LLMs ranging from 8B to 236B parameters across prompt strategy and temperature configurations. The highest-performing LLM achieved a micro-averaged F1 score of 0.22, establishing the empirical baseline for multi-label ATT&CK classification on complex unstructured CTI. Parameter size showed a statistically significant positive correlation with F1 score. Prompt strategy and temperature produced no statistically significant gains across model configurations. These results indicate that current open-source LLMs are insufficient for production-grade ATT&CK classification. The dataset, benchmark, and findings provide a reproducible foundation for future CTI research.

威胁情报大模型评估安全分析多标签分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。