arXiv:2510.11974cs.CRcs.AI2025-10KDD被引 3

构建首个异构威胁情报检索增强评估基准,揭示LLM在实战中的关键瓶颈。

CTIConnect: A Benchmark for Retrieval-Augmented LLMs over Heterogeneous Cyber Threat Intelligence

  • 整合五类异构威胁情报源,构建1860个专家验证的问答对
  • 发现跨源语义鸿沟因任务类型不同需差异化检索策略
  • 适合安全研究者与AI系统设计者参考,指导实际部署

网络威胁情报(CTI)是现代网络安全的核心,使组织能够主动应对不断演变的威胁。然而,结构化知识库(如CVE、CWE、CAPEC、MITRE ATT&CK)和非结构化威胁报告的海量异构数据远超人工分析能力。大语言模型(LLMs)强大的上下文理解与推理能力正推动其在CTI任务中的应用。但现有评估缺乏在检索增强场景下、支持真实分析师依赖的多源知识访问的评测框架。为此,我们提出CTIConnect,一个系统性评估检索增强型LLMs的基准。该基准整合五类异构CTI源,构建1,860个专家验证的问答对,覆盖九项任务,分属实体链接、多文档合成与实体归因三类。对十种先进LLMs的实验表明,跨源语义差距在不同任务类别中表现各异,需采用根本不同的检索策略;性能瓶颈在检索基础设施与证据利用之间随任务变化。领域特定策略优于更强的通用检索范式(retrieve-then-rerank、IRCoT),说明弥合差距需结构性干预而非通用优化。这些发现贯穿全部十种模型,在完整基准上稳定,并在2008–2025年时间跨度下保持一致,为异构CTI生态系统的可扩展检索架构设计提供可操作指导。

原文摘要 · Abstract (English)

Cyber Threat Intelligence (CTI) is foundational to modern cybersecurity, enabling organizations to proactively defend against evolving threats. However, the sheer volume and heterogeneity of CTI data, spanning structured knowledge bases (CVE, CWE, CAPEC, MITRE ATT&CK) and unstructured threat reports, far exceed the capacity of manual analysis. The strong contextual understanding and reasoning of Large Language Models (LLMs) have driven growing interest in applying them to CTI tasks. Yet no existing benchmark evaluates LLMs in a retrieval-augmented setting with a proper evaluation harness that grants access to the heterogeneous domain knowledge sources analysts rely on in practice. To address this gap, we present CTIConnect, a benchmark for systematically evaluating retrieval-augmented LLMs across the CTI task landscape. We construct a unified evaluation environment integrating five heterogeneous CTI sources into 1,860 expert-verified QA pairs spanning nine tasks across three categories: Entity Linking, Multi-Document Synthesis, and Entity Attribution. Extensive experiments on ten state-of-the-art LLMs reveal that the cross-source semantic gap manifests differently across task categories, demanding fundamentally different retrieval strategies, and that the performance bottleneck shifts between retrieval infrastructure and evidence utilization depending on the task. Our domain-specific strategies further outperform stronger general-purpose retrieval paradigms (retrieve-then-rerank, IRCoT), showing that closing this gap requires structural interventions rather than generic retrieval improvements. These findings hold across all ten LLMs, remain consistent on the full benchmark, and stay stable under temporal splits spanning 2008-2025. Together, they provide actionable guidance for designing scalable retrieval architectures over heterogeneous CTI ecosystems.

威胁情报检索增强LLM评估安全AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。