arXiv:2605.25836cs.CRcs.AI2026-05被引 2

通过分步验证提升威胁情报中的攻击技术提取精度与召回率

TTPrint: Evidence-Grounded TTP Extraction via Diverge-then-Converge Verification

论文配图:TTPrint: Evidence-Grounded TTP Extraction via Diverge-then-Converge Verification
图 1 · 摘自论文原文
  • 先广撒网提取候选技术,再通过证据定位严格验证
  • 在两个新数据集上分别取得87.39%和76.48%的宏平均F1
  • 适合需要高可靠性的安全分析人员和自动化威胁研判系统

从网络威胁情报报告中提取MITRE ATT&CK技术是一项开放集、多标签任务,要求高召回率(不遗漏)与高精确率(不误报)。现有方法——规则、监督学习和大模型——难以兼顾:规则与监督方法泛化能力差,而将候选生成与验证合并的单步大模型方法同时受限于召回率与精确率。我们提出TTPrint,其设计受人类分析师工作方式启发:先发散式提取,再收敛式验证。在发散阶段,报告被分解为原子行为并广泛提出候选技术;确定性片段定位阶段将每个候选锚定到原文中的具体证据窗口;收敛验证阶段仅保留同时满足局部证据与权威MITRE定义的技术。我们贡献了两个评估资源——清理后的TRAM基准(TRAM-Clean)和新的标注数据集(TTPrint-Bench),以解决现有基准的标注噪声问题,并推动任务向文档级提取演进。在TRAM-Clean和TTPrint-Bench上,TTPrint分别取得76.48%和87.39%的宏平均F1,优于领先基线63.5%和29.4%。六种大模型的多骨干分析及阈值敏感性研究进一步验证了其跨模型泛化能力,并提供实用参数选择指导。

原文摘要 · Abstract (English)

Extracting MITRE ATT&CK techniques from cyber threat intelligence (CTI) reports is an open-set, multi-label problem requiring both high recall (not missing techniques) and high precision (not hallucinating unsupported ones). Existing methods--rule-based, supervised, and LLM-based--struggle to achieve both: rule-based and supervised approaches lack generalizability across diverse attack descriptions, while LLM-based approaches that couple candidate generation and validation within a single inference step suffer from limited recall and precision simultaneously. We propose TTPrint, which addresses this challenge through a diverge-then-converge design inspired by how human analysts work: first extracting broadly, then verifying rigorously. In the divergent phase, reports are decomposed into atomic behaviors and candidate techniques are proposed broadly. A deterministic span localization stage then anchors each candidate to a specific evidence window in the source text. A convergent verification stage retains only candidates supported by both the localized evidence and the authoritative MITRE definition. We contribute two evaluation resources--a cleaned TRAM benchmark (TRAM-Clean) and a new annotated dataset (TTPrint-Bench)--to address known annotation noise in existing benchmarks and elevate the task to document-level TTP extraction. On TRAM-Clean and TTPrint-Bench, TTPrint achieves 76.48% and 87.39% macro-F1 respectively, outperforming the leading baseline by 63.5% and 29.4%. A multi-backbone analysis across six LLMs and a threshold sensitivity study further demonstrate generalizability across model choices and provide practical guidance for parameter selection.

威胁情报信息提取大模型应用安全分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。