arXiv:2507.16852cs.CRcs.AI2025-07

用大模型生成虚假威胁情报,解决标注数据少的问题

SynthCTI: LLM-Driven Synthetic CTI Generation to enhance MITRE Technique Mapping

  • 用聚类提取语义上下文,指导大模型生成逼真威胁描述
  • 使小模型分类准确率提升48.6%,超过未增强的大模型
  • 适合安全研究者和自动化威胁分析系统使用

网络威胁情报(CTI)挖掘需从非结构化数据中提取结构化信息,以理解攻击者行为。将威胁描述映射到MITRE ATT&CK技术是关键任务,但常依赖人工,耗时且需专家知识。自动方法面临高质量标注数据稀缺和类别不平衡问题。尽管专用大模型如SecureBERT表现更好,但现有研究多关注模型架构而非数据瓶颈。本文提出SynthCTI,一种数据增强框架,用于为低频的MITRE ATT&CK技术生成高质量合成威胁语句。该方法基于聚类提取训练数据的语义上下文,引导大模型生成在词汇上多样、语义上忠实的合成句子。我们在两个公开数据集CTI-to-MITRE和TRAM上评估,采用不同规模的LLM。引入合成数据后,宏观F1得分显著提升:ALBERT从0.35升至0.52(相对提升48.6%),SecureBERT达0.6558(较0.4412提升)。值得注意的是,经合成数据增强的小模型性能优于未增强的大模型,证明数据生成对构建高效、精准的威胁分类系统具有关键价值。

原文摘要 · Abstract (English)

Cyber Threat Intelligence (CTI) mining involves extracting structured insights from unstructured threat data, enabling organizations to understand and respond to evolving adversarial behavior. A key task in CTI mining is mapping threat descriptions to MITRE ATT\&CK techniques. However, this process is often performed manually, requiring expert knowledge and substantial effort. Automated approaches face two major challenges: the scarcity of high-quality labeled CTI data and class imbalance, where many techniques have very few examples. While domain-specific Large Language Models (LLMs) such as SecureBERT have shown improved performance, most recent work focuses on model architecture rather than addressing the data limitations. In this work, we present SynthCTI, a data augmentation framework designed to generate high-quality synthetic CTI sentences for underrepresented MITRE ATT\&CK techniques. Our method uses a clustering-based strategy to extract semantic context from training data and guide an LLM in producing synthetic CTI sentences that are lexically diverse and semantically faithful. We evaluate SynthCTI on two publicly available CTI datasets, CTI-to-MITRE and TRAM, using LLMs with different capacity. Incorporating synthetic data leads to consistent macro-F1 improvements: for example, ALBERT improves from 0.35 to 0.52 (a relative gain of 48.6\%), and SecureBERT reaches 0.6558 (up from 0.4412). Notably, smaller models augmented with SynthCTI outperform larger models trained without augmentation, demonstrating the value of data generation methods for building efficient and effective CTI classification systems.

威胁情报数据增强大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。