arXiv:2511.01144cs.CRcs.AI2025-11被引 7

构建动态评测基准,评估大模型在网络安全情报中的表现

AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence

  • 设计新数据生成流程与去重机制,提升评测质量
  • 发现主流大模型在威胁溯源等任务上表现仍不足
  • 适合关注网络安全自动化与模型评测的研究者

大型语言模型(LLMs)在自然语言推理方面表现出色,但在网络安全情报(CTI)领域的应用仍受限。CTI分析需从海量非结构化报告中提炼可操作知识,正是LLMs可大幅降低分析师工作量的场景。现有CTIBench已涵盖多项CTI任务的评测。本文在此基础上提出AthenaBench,包含改进的数据集构建流程、去重机制、优化的评估指标,以及新增的风险缓解策略生成任务。我们评估了12个LLM,包括GPT-5、Gemini-2.5 Pro等前沿闭源模型,以及来自LLaMA和Qwen系列的7个开源模型。尽管闭源模型整体表现更优,但在威胁主体归因、风险缓解等推理密集型任务上仍表现不佳,开源模型差距更大。结果揭示当前大模型在复杂推理能力上的根本局限,凸显开发专用于CTI工作流与自动化的模型的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated strong capabilities in natural language reasoning, yet their application to Cyber Threat Intelligence (CTI) remains limited. CTI analysis involves distilling large volumes of unstructured reports into actionable knowledge, a process where LLMs could substantially reduce analyst workload. CTIBench introduced a comprehensive benchmark for evaluating LLMs across multiple CTI tasks. In this work, we extend CTIBench by developing AthenaBench, an enhanced benchmark that includes an improved dataset creation pipeline, duplicate removal, refined evaluation metrics, and a new task focused on risk mitigation strategies. We evaluate twelve LLMs, including state-of-the-art proprietary models such as GPT-5 and Gemini-2.5 Pro, alongside seven open-source models from the LLaMA and Qwen families. While proprietary LLMs achieve stronger results overall, their performance remains subpar on reasoning-intensive tasks, such as threat actor attribution and risk mitigation, with open-source models trailing even further behind. These findings highlight fundamental limitations in the reasoning capabilities of current LLMs and underscore the need for models explicitly tailored to CTI workflows and automation.

大模型评测网络安全推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。