arXiv:2410.09114cs.CRcs.AI2024-10被引 17

评测大模型在真实网络攻击中的能力,发现顶级模型能攻陷多种系统。

Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities

  • 构建3CB基准测试框架,专评大模型的实战化进攻能力。
  • GPT-4o和Claude 3.5 Sonnet可完成二进制分析到网页攻击的多类任务。
  • 开源小模型能力有限,为监管和安全部署提供关键评估工具。

大语言模型代理有潜力革新网络安全防御,但其进攻能力尚未被充分理解。为应对新兴威胁,模型开发者与政府正评估基础模型的网络能力,但现有评估缺乏透明度且未聚焦进攻面。为此,我们提出灾难性网络能力基准(3CB),一个旨在严格评估大模型代理真实世界进攻能力的新框架。对现代大模型在3CB上的评估显示,前沿模型如GPT-4o和Claude 3.5 Sonnet可在二进制分析、网页技术等多个领域执行侦察与利用等进攻任务;而较小的开源模型则表现出明显局限。我们的软件解决方案与对应基准,有助于缩小快速提升的能力与稳健评估之间的差距,助力这些强大技术的安全部署与监管。

原文摘要 · Abstract (English)

LLM agents have the potential to revolutionize defensive cyber operations, but their offensive capabilities are not yet fully understood. To prepare for emerging threats, model developers and governments are evaluating the cyber capabilities of foundation models. However, these assessments often lack transparency and a comprehensive focus on offensive capabilities. In response, we introduce the Catastrophic Cyber Capabilities Benchmark (3CB), a novel framework designed to rigorously assess the real-world offensive capabilities of LLM agents. Our evaluation of modern LLMs on 3CB reveals that frontier models, such as GPT-4o and Claude 3.5 Sonnet, can perform offensive tasks such as reconnaissance and exploitation across domains ranging from binary analysis to web technologies. Conversely, smaller open-source models exhibit limited offensive capabilities. Our software solution and the corresponding benchmark provides a critical tool to reduce the gap between rapidly improving capabilities and robustness of cyber offense evaluations, aiding in the safer deployment and regulation of these powerful technologies.

大模型安全攻防评测3CB基准网络攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。