SIR-Bench评估安全响应智能体的深度调查能力,区分真实取证与简单复述。
SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents
- 构建真实云环境回放框架OUAT,生成可验证的调查数据
- 三类指标评估:误报率73.4%,每案例发现5.67个新证据
- 用对抗性大模型裁判,要求提供具体证据才计分
我们提出SIR-Bench,一个包含794个测试用例的基准,用于评估自主安全事件响应智能体,能够区分真正的取证调查与仅复述告警。该基准源自129个匿名化事件模式,并经专家验证真值。SIR-Bench不仅衡量智能体是否做出正确分类决策,还评估其通过主动调查发现新证据的能力。为构建该基准,我们开发了Once Upon A Threat(OUAT)框架,在受控云环境中重演真实事件模式,生成具有可测量调查结果的真实遥测数据。评估方法引入三个互补指标:分类准确率(M1)、新发现证据能力(M2)和工具使用恰当性(M3),通过对抗性LLM作为裁判,将举证责任反转——必须提供具体取证证据才能获得分数。在该基准上评估我们的SIR智能体,结果显示97.1%的真正例检测率、73.4%的假正例拒识率,以及每案例平均5.67个新关键发现,为未来调查智能体提供了可量化的基准。
原文摘要 · Abstract (English)
We present SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents that distinguishes genuine forensic investigation from alert parroting. Derived from 129 anonymized incident patterns with expert-validated ground truth, SIR-Bench measures not only whether agents reach correct triage decisions, but whether they discover novel evidence through active investigation. To construct SIR-Bench, we develop Once Upon A Threat (OUAT), a framework that replays real incident patterns in controlled cloud environments, producing authentic telemetry with measurable investigation outcomes. Our evaluation methodology introduces three complementary metrics: triage accuracy (M1), novel finding discovery (M2), and tool usage appropriateness (M3), assessed through an adversarial LLM-as-Judge that inverts the burden of proof -- requiring concrete forensic evidence to credit investigations. Evaluating our SIR agent on the benchmark demonstrates 97.1% true positive (TP) detection, 73.4% false positive (FP) rejection, and 5.67 novel key findings per case, establishing a baseline against which future investigation agents can be measured.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。