arXiv:2510.01654cs.CLcs.AI2025-10被引 2

为安全智能体设计评估框架,量化其闭环防御能力。

SoK: Measuring What Matters for Closed-Loop Security Agents

  • 构建闭环安全智能体能力映射框架,连接安全流程与智能体能力。
  • 提出CLC评分,综合衡量闭环程度与实际效果。
  • 适用于研究者与工程师评估和改进自动化安全系统。

网络安全是一场永不停歇的攻防对抗,以AI驱动的攻击手段演进速度远超传统防御响应能力。当前研究与工具分散于孤立的防御功能中,形成被攻击者利用的盲区。能够整合探测、漏洞验证、修复与验证于一体的自主闭环智能体具有潜力,但该领域缺乏三大基础:定义安全生命周期内智能体能力的框架、评估闭环智能体的系统方法,以及实践中的性能基准。本文提出CLASP框架,将安全生命周期(侦察、利用、根因分析、补丁生成、验证)与核心智能体能力(规划、工具使用、记忆、推理、反思与感知)对齐,提供统一的评估术语与标准。通过对21个代表性工作的应用,我们揭示了各系统的优势与能力缺口。进而定义闭环能力(CLC)得分,一个综合量化闭环程度与运行效能的复合指标,并明确闭环基准的构建要求。CLASP与CLC得分共同提供了推进智能体功能性能与测量闭环安全能力所需的词汇、诊断工具与量化手段。

原文摘要 · Abstract (English)

Cybersecurity is a relentless arms race, with AI driven offensive systems evolving faster than traditional defenses can adapt. Research and tooling remain fragmented across isolated defensive functions, creating blind spots that adversaries exploit. Autonomous agents capable of integrating, exploit confirmation, remediation, and validation into a single closed loop offer promise, but the field lacks three essentials: a framework defining the agentic capabilities of security systems across security life cycle, a principled method for evaluating closed loop agents, and a benchmark for measuring their performance in practice. We introduce CLASP: the Closed-Loop Autonomous Security Performance framework which aligns the security lifecycle (reconnaissance, exploitation, root cause analysis, patch synthesis, validation) with core agentic capabilities (planning, tool use, memory, reasoning, reflection & perception) providing a common vocabulary and rubric for assessing agentic capabilities in security tasks. By applying CLASP to 21 representative works, we map where systems demonstrate strengths, and where capability gaps persist. We then define the Closed-Loop Capability (CLC) Score, a composite metric quantifying both degree of loop closure and operational effectiveness, and outline the requirements for a closed loop benchmark. Together, CLASP and the CLC Score, provide the vocabulary, diagnostics, and measurements needed to advance both function level performance and measure closed loop security agents.

安全智能体闭环评估框架设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。