诊断式评估大模型在攻击链重建中的推理错误,揭示不同阶段的短板。
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

- 构建分阶段评估框架,追踪从证据检索到链路排序的每一步推理
- 69个场景测试显示最强模型仅正确还原39.6%的参考步骤
- 小模型难整合证据,大模型则卡在证据排序环节
大型语言模型(LLM)代理通过检索和解析异构遥测数据,有望实现攻击链重建。然而现有基准主要评估最终输出或整体准确率,难以揭示错误在中间推理阶段的产生与传播。本文提出DiagChain,一个面向证据锚定攻击链重建的诊断性基准,支持分阶段评估。DiagChain包含MAIN-69,涵盖多种操作系统、证据噪声水平和链长的69个场景。同时引入基于证据中心的检索增强生成(ECRAG),将证据检索与重构链的动态结构表示耦合。设计五项互补指标,用于评估重建过程的不同阶段,支持系统性故障诊断。对6种LLM的评估显示,即使最强配置在MAIN-69的849个参考步骤中也仅成功39.6%。分析表明,小型模型在将检索证据纳入输出这一基础任务上表现不佳,而大型模型虽能推进至后续阶段,但正确排序证据成为主要瓶颈。结果验证了超越端到端准确率的诊断评估的重要性,并为提升证据锚定的网络安全代理提供可操作洞察。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions. However, existing benchmarks mainly evaluate final outputs or aggregate accuracy, providing limited insight into how errors arise and propagate across intermediate reasoning stages. We present DiagChain, a diagnostic benchmark for evidence-grounded attack chain reconstruction that enables stage-wise evaluation of LLM agents. DiagChain includes MAIN-69, a suite of 69 scenarios spanning multiple operating systems, evidence noise levels, and chain lengths. It further introduces Evidence-Centric Retrieval-Augmented Generation (ECRAG), which couples evidence retrieval with an evolving structured representation of the reconstructed chain. Five complementary metrics are introduced to assess distinct stages of the reconstruction process and support systematic failure diagnosis. Based on evaluations using 6 LLMs, DiagChain reveals that even the strongest configuration succeeds on only 39.6% of the 849 reference steps in MAIN-69. Our analysis further shows that smaller models struggle with the more basic task of incorporating retrieved evidence into their outputs, whereas larger models can proceed to later steps, where correctly ordering that evidence becomes the main bottleneck. These results validate the importance of diagnostic evaluation beyond end-to-end accuracy and provide actionable insights for improving evidence-grounded cybersecurity agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。