arXiv:2505.18384cs.CRcs.AI2025-05NeurIPS被引 5

实验证明小算力下黑客可迭代提升攻击代理能力超40%。

Dynamic Risk Assessments for Offensive Cybersecurity Agents

  • 在固定算力下模拟真实对抗,评估攻击代理的动态进化能力。
  • 仅用8个H100 GPU小时,攻击能力提升超过40%。
  • 提醒安全评估需考虑对手的持续优化能力,适合攻防研究者。

基础模型正日益具备自主编程能力,可能被用于自动化危险的网络攻击。当前前沿模型审计虽能检测安全风险,但大多未考虑真实世界中攻击者的自由度。尤其在强验证机制与经济激励下,攻击代理可通过迭代改进被优化。本文主张在网络安全评估中引入更广义威胁模型,强调状态与非状态环境下攻击者可用的自由度,同时限定计算预算。实验表明,在固定算力(本研究中为8 H100 GPU小时)下,攻击者可使代理在InterCode CTF上的能力相较基线提升超过40%,且无需外部协助。这一结果凸显动态风险评估的重要性,能更真实反映潜在威胁。

原文摘要 · Abstract (English)

Foundation models are increasingly becoming better autonomous programmers, raising the prospect that they could also automate dangerous offensive cyber-operations. Current frontier model audits probe the cybersecurity risks of such agents, but most fail to account for the degrees of freedom available to adversaries in the real world. In particular, with strong verifiers and financial incentives, agents for offensive cybersecurity are amenable to iterative improvement by would-be adversaries. We argue that assessments should take into account an expanded threat model in the context of cybersecurity, emphasizing the varying degrees of freedom that an adversary may possess in stateful and non-stateful environments within a fixed compute budget. We show that even with a relatively small compute budget (8 H100 GPU Hours in our study), adversaries can improve an agent's cybersecurity capability on InterCode CTF by more than 40\% relative to the baseline -- without any external assistance. These results highlight the need to evaluate agents' cybersecurity risk in a dynamic manner, painting a more representative picture of risk.

网络安全攻击代理动态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。