arXiv:2607.15263cs.CRcs.AI2026-07

评估安全模型时,不仅要看成功率,还要看成本效率。

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

论文配图:Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
图 1 · 摘自论文原文
  • 按固定成本比较红蓝队任务表现,分解推理与工具开销
  • 进攻任务随计算资源增加而提升,开源模型可逼近闭源系统
  • 防守任务更依赖工具使用策略,而非单纯算力

当前安全智能体评估多聚焦于宽松推理预算下的最大进攻能力,强调漏洞发现、攻击开发、渗透测试和攻防竞赛完成度。这类指标虽有用但不完整:实际安全操作中,每一步推理、工具调用、遥测查询和信息丰富请求均消耗预算。本文在进攻型Cybench挑战和防御型Splunk BOTS v1调查任务上,通过成本-成功视角评估语言模型安全智能体。不再仅报告最优成功率,而是固定成本水平下对比模型表现,并分解性能为推理开销与工具开销。结果表明红队与蓝队任务存在显著不同的扩展规律:进攻型CTF性能随测试时计算资源增加而提升,规模化开源模型可接近前沿闭源系统且保持成本竞争力;而防守型SOC调查并不呈现相同扩展性:成功更依赖工具使用的纪律性、遥测导航能力及选择性信息丰富,而非单纯推理预算。我们主张安全智能体基准应同时衡量经济效率与任务适配性。成本感知、以SOC为核心的评估能更清晰揭示当前哪些模型真正实用,以及防御型智能体仍需改进之处。相关结果已发布于交互式网站:https://evals.frontier.security。

原文摘要 · Abstract (English)

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.

安全智能体成本评估攻防对抗蓝队效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。