arXiv:2509.23571cs.CRcs.AI2025-09被引 6

用标准化流程提升大模型在蓝队威胁狩猎中的实战能力

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

  • 构建分步模块化框架,将威胁狩猎拆解为30个可操作任务
  • 相比自由推理,标准化流程使大模型准确率显著提升
  • 适合安全研究者和蓝队人员评估大模型实战效能

随着网络威胁规模与复杂性持续增长,蓝队防御者亟需先进工具主动发现并应对风险。大语言模型(LLMs)在增强威胁分析方面展现出潜力,但其在真实蓝队威胁狩猎场景中的有效性尚未充分探索。本文提出CyberTeam基准,旨在指导LLMs开展蓝队实践。CyberTeam采用两阶段标准化工作流:首先建模真实威胁狩猎流程,捕捉从威胁溯源到事件响应各分析任务间的依赖关系;其次,针对每项任务配置特定操作模块,将威胁狩猎转化为一系列有据可依的推理步骤,每个步骤对应具体操作,并按任务依赖顺序排列。在此框架下,LLMs通过模块化步骤执行威胁狩猎任务。整体上,CyberTeam集成30个任务与9个操作模块,引导LLMs完成标准化威胁分析。我们评估了主流LLMs及先进网络安全代理,对比其在标准流程与开放式推理策略下的表现。结果表明,标准化设计带来显著性能提升,同时揭示开放式推理在真实威胁狩猎中的局限性。

原文摘要 · Abstract (English)

As cyber threats continue to grow in scale and sophistication, blue team defenders increasingly require advanced tools to proactively detect and mitigate risks. Large Language Models (LLMs) offer promising capabilities for enhancing threat analysis. However, their effectiveness in real-world blue team threat-hunting scenarios remains insufficiently explored. This paper presents CyberTeam, a benchmark designed to guide LLMs in blue teaming practice. CyberTeam constructs a standardized workflow in two stages. First, it models realistic threat-hunting workflows by capturing the dependencies among analytical tasks from threat attribution to incident response. Next, each task is addressed through a set of operational modules tailored to its specific analytical requirements. This transforms threat hunting into a structured sequence of reasoning steps, with each step grounded in a discrete operation and ordered according to task-specific dependencies. Guided by this framework, LLMs are directed to perform threat-hunting tasks through modularized steps. Overall, CyberTeam integrates 30 tasks and 9 operational modules to guide LLMs through standardized threat analysis. We evaluate both leading LLMs and state-of-the-art cybersecurity agents, comparing CyberTeam against open-ended reasoning strategies. Our results highlight the improvements enabled by standardized design, while also revealing the limitations of open-ended reasoning in real-world threat hunting.

大模型安全威胁狩猎蓝队基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。