arXiv:2608.28394cs.CRcs.CL2026-08

用攻击行为锚定跨源威胁情报,实现更精准的知识图谱构建。

BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence

论文配图:BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence
图 1 · 摘自论文原文
  • 以MITRE ATT&CK行为为锚点,统一不同报告中的威胁实体
  • 在8,395个元素上比基线高23%,3,487个跨源对齐项上高9%
  • 适合网络安全分析与自动化威胁情报整合场景

网络威胁情报(CTI)是现代网络安全防御的基础,但其大部分内容存在于非结构化报告中,数量庞大且异构性高,远超人工分析能力,推动了从CTI报告自动生成知识图谱的研究。然而,现有方法主要局限于单份报告内提取部分信息,未探索跨源场景——同一威胁在不同报告中使用不同名称。本文关键洞察是:一旦将攻击行为映射到标准的MITRE ATT&CK(攻击技术目录),即可作为锚点统一报告间信息。攻击行为是报告描述的恶意行动,而上下文实体(如威胁组织、活动、受影响产品)和指标(如IP地址)则是其参与者与痕迹。将它们关联到这些锚点,可将每份报告的图谱置于同一规范空间。我们提出BEACON,一种基于大模型的跨源CTI知识图谱构建框架。第一阶段采用先提出后验证范式,将候选实体锚定于报告证据与官方ATT&CK定义,抑制大模型误判与幻觉;第二阶段通过分层对齐策略合并图谱,按确定性递减顺序应用信号:字符级相似性、语义相似性、重叠技术邻域,迭代融合邻域信息。目前无公开基准能链接实体至技术锚点或提供跨源对齐真值。因此,我们从34个来源构建并发布两个人工标注数据集:据知最大报告级CTI抽取数据集(8,395个元素)和首个跨源整合数据集(3,487个对齐项)。在上述数据集上,BEACON性能优于所有基线至少23%和9%。

原文摘要 · Abstract (English)

Cyber threat intelligence (CTI) is foundational to modern cyber defense, yet much of it resides in unstructured reports whose volume and heterogeneity far exceed manual analysis, motivating research on automatically constructing knowledge graphs from CTI reports. However, existing approaches mainly extract partial information within a single report, leaving the cross-source setting unexplored, where the same threat is given unrelated names. Our key insight is that attack behaviors, once mapped to MITRE ATT&CK (a standardized catalog of attack techniques), can anchor the rest of a report. Attack behaviors are the adversarial actions a report describes, while contextual entities (e.g., threat actors, campaigns, and affected products) and Indicators of Compromise (IoCs; e.g., IP addresses) are their participants and traces. Attaching them to these anchors places every per-report graph in one canonical space. We realize this insight in BEACON, an LLM-driven framework for cross-source CTI knowledge graph construction. Its first stage extracts each report into a graph under a propose-then-verify paradigm, grounding candidates in report evidence and official ATT&CK definitions, to suppress LLM misclassification and hallucination. Its second stage merges these graphs with a hierarchical alignment strategy that applies signals in decreasing order of determinism, from character-level and semantic similarity to overlapping technique neighborhoods, iterating as merges pool neighborhoods. No existing benchmark links entities to technique anchors or provides cross-source alignment ground truth. We therefore construct and release two human-annotated datasets from 34 sources: to our knowledge the largest for report-level CTI extraction (8,395 elements) and the first for cross-source consolidation (3,487). On them, BEACON outperforms all baselines by at least 23% and 9%, respectively.

威胁情报知识图谱LLMATT&CK

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。