构建首个面向恶意软件分析与威胁情报推理的开源评测基准。
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
- 设计针对恶意软件分析和威胁情报推理的双任务评测体系。
- 大模型表现优于小模型,但推理能力提升有限。
- 适合安全AI开发者与红蓝对抗研究者使用。
当前网络安全防御面临海量安全告警、威胁情报信号及业务环境变化,亟需AI系统辅助运营。尽管大语言模型(LLMs)有望自动化与规模化安全运维,但现有评估未能覆盖真实场景。这导致开发者缺乏明确方向,用户难以选择有效模型。与此同时,攻击方已开始利用AI扩大攻击规模,凸显开源评测基准的重要性。为此,我们提出CyberSOCEval,作为CyberSecEval 4中的开源评测套件,涵盖恶意软件分析与威胁情报推理两大核心任务。实验表明,更大更现代的LLMs表现更优,验证了训练规模定律;但采用测试时扩展的推理模型在安全分析任务中未获类似编程与数学任务中的性能提升,说明其未充分学习网络安全推理能力。当前模型远未达到评测上限,证明CyberSOCEval对安全AI发展构成实质性挑战。
原文摘要 · Abstract (English)
Today's cyber defenders are overwhelmed by a deluge of security alerts, threat intelligence signals, and shifting business context, creating an urgent need for AI systems to enhance operational security work. While Large Language Models (LLMs) have the potential to automate and scale Security Operations Center (SOC) operations, existing evaluations do not fully assess the scenarios most relevant to real-world defenders. This lack of informed evaluation impacts both AI developers and those applying LLMs to SOC automation. Without clear insight into LLM performance in real-world security scenarios, developers lack a north star for development, and users cannot reliably select the most effective models. Meanwhile, malicious actors are using AI to scale cyber attacks, highlighting the need for open source benchmarks to drive adoption and community-driven improvement among defenders and model developers. To address this, we introduce CyberSOCEval, a new suite of open source benchmarks within CyberSecEval 4. CyberSOCEval includes benchmarks tailored to evaluate LLMs in two tasks: Malware Analysis and Threat Intelligence Reasoning--core defensive domains with inadequate coverage in current benchmarks. Our evaluations show that larger, more modern LLMs tend to perform better, confirming the training scaling laws paradigm. We also find that reasoning models leveraging test time scaling do not achieve the same boost as in coding and math, suggesting these models have not been trained to reason about cybersecurity analysis, and pointing to a key opportunity for improvement. Finally, current LLMs are far from saturating our evaluations, showing that CyberSOCEval presents a significant challenge for AI developers to improve cyber defense capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。