构建大规模漏洞评估平台,测试AI在真实场景中发现和复现漏洞的能力。
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
- 基于188个软件项目中的1507个真实漏洞,动态评估AI生成漏洞复现代码的能力。
- 顶尖模型成功率仅约20%,体现当前AI在网络安全任务中的显著局限性。
- 不仅可评测模型性能,还能发现34个零日漏洞,适合安全研究与AI评估团队使用。
AI代理在重塑网络安全方面具有巨大潜力,因此对其能力进行全面评估至关重要。然而,现有评估方法存在局限:基于小规模基准,仅衡量静态结果,无法捕捉真实世界安全挑战的全貌。为此,我们提出CyberGym,一个包含188个软件项目中1507个真实漏洞的大规模基准。该平台可根据不同漏洞分析场景灵活调整,主要任务是让代理在仅有漏洞文本描述和对应代码库的情况下,生成可复现漏洞的原型测试用例。大规模评估表明,CyberGym能有效区分不同代理与模型的网络安全能力。即使表现最优的组合,成功率也仅为约20%,凸显了该基准的整体难度。此外,除静态评估外,我们在实验中成功发现了34个零日漏洞及18个历史遗留不完整补丁。这些结果表明,CyberGym不仅是衡量AI在网络安全领域进展的稳健基准,更是一个可直接产生真实世界安全影响的平台。
原文摘要 · Abstract (English)
AI agents have significant potential to reshape cybersecurity, making a thorough assessment of their capabilities critical. However, existing evaluations fall short, because they are based on small-scale benchmarks and only measure static outcomes, failing to capture the full, dynamic range of real-world security challenges. To address these limitations, we introduce CyberGym, a large-scale benchmark featuring 1,507 real-world vulnerabilities across 188 software projects. Adjustable to different vulnerability analysis settings, CyberGym primarily tasks agents with generating a proof-of-concept test that reproduces a vulnerability, given only its text description and the corresponding codebase. Our extensive evaluation highlights that CyberGym effectively differentiates agents' and models' cybersecurity capabilities. Even the top-performing combinations only achieve a ~20% success rate, demonstrating the overall difficulty of CyberGym. Beyond static benchmarking, we show that CyberGym leads to the discovery of 34 zero-day vulnerabilities and 18 historically incomplete patches. These results underscore that CyberGym is not only a robust benchmark for measuring AI's progress in cybersecurity but also a platform for creating direct, real-world security impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。