用小模型团队构建更稳定精准的网络安全知识图谱
TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction

- 拆解任务为提取、类型识别、验证等专业小模型代理
- 3B-8B轻量模型实现更高召回率与图结构一致性
- 适合需要可控、低成本部署的网络安全分析场景
网络安全威胁情报报告多为非结构化、异构且含噪声,难以直接用于自动化分析。网络安全知识图谱(CSKG)虽能结构化呈现攻击者实体、行为与关系,但从自由文本中构建仍具挑战。现有方法依赖庞大单一语言模型端到端抽取与补全,导致成本高、控制难、性能不稳定。本文提出TACTIC-KG,一种基于代理的框架,将任务分解为提取、类型标注、验证与整理等模块化、专业化的小模型代理。采用3B至8B参数的轻量模型,在提升稳定性、召回率与图一致性的前提下降低部署成本。在人工标注的威胁情报报告上评估显示,该方法在抽取F1值、类型准确率及图结构相似性上均优于主流单体上下文学习基线。
原文摘要 · Abstract (English)
Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for automated analysis and reasoning. Cybersecurity Knowledge Graphs (CSKGs) provide a structured representation of adversarial entities, actions, and relations, but constructing such graphs from free-text CTI remains a challenge. Recent approaches rely on monolithic Large Language Models (LLMs) to perform end-to-end extraction and completion, leading to high cost, limited controllability, and unstable performance. This paper introduces TACTIC-KG, an agentic framework for CSKG construction that decomposes the task into modular, specialized LLM agents responsible for extraction, typing, verification, and curation. Using lightweight models (3B--8B), TACTIC-KG improves stability, recall, and graph consistency while reducing deployment cost. We implement and evaluate TACTIC-KG against recent state-of-the-art systems. Experiments on human-annotated CTI reports show that agent specialization consistently outperforms larger monolithic in-context-learning (ICL) baselines in extraction F1-score, typing accuracy, and structural graph similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。