构建真实漏洞攻击评估基准,测试大模型代理的实战攻防能力
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
- 基于真实高危漏洞构建沙箱环境,模拟现实攻击场景
- 顶尖代理仅能成功利用13%的漏洞,暴露当前智能攻防短板
- 适合安全研究者、红队演练人员及防御系统开发者参考
大型语言模型(LLM)代理正展现出自主发起网络攻击的能力,对现有应用构成严重威胁。这一风险凸显了建立真实世界基准以评估LLM代理利用网页应用漏洞能力的紧迫性。然而,现有基准多局限于抽象化的夺旗竞赛,或缺乏全面覆盖。构建真实漏洞基准需专业技能复现攻击,并系统化评估不可预测威胁。为此,我们提出CVE-Bench,一个基于关键严重性通用漏洞披露(CVE)的真实网络安全基准。在该基准中,我们设计了一个沙箱框架,使LLM代理能够在贴近真实条件的场景下攻击存在漏洞的Web应用,同时提供有效的攻击评估。我们的评估显示,当前最先进的代理框架最多可解决13%的漏洞。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the ability of LLM agents to exploit web application vulnerabilities. However, existing benchmarks fall short as they are limited to abstracted Capture the Flag competitions or lack comprehensive coverage. Building a benchmark for real-world vulnerabilities involves both specialized expertise to reproduce exploits and a systematic approach to evaluating unpredictable threats. To address this challenge, we introduce CVE-Bench, a real-world cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Our evaluation shows that the state-of-the-art agent framework can resolve up to 13% of vulnerabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。