构建可扩展的端到端网络安全评估基准,测试AI自主发现修复漏洞能力
CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

- 基于开源漏洞数据自动构建真实环境,实现从漏洞发现到修复的全流程评测
- 涵盖139个开源项目中的920个真实漏洞,支持大规模自动化评估
- 适合研究自主网络安全智能体的学者和开发者使用
AI有望通过实现自主检测、分析和修复软件漏洞来变革网络安全。然而,现有AI系统在网络安全方面的评估规模或范围有限,无法涵盖真实世界中漏洞发现与修复的完整生命周期。为弥补这一空白,我们提出CyberGym-E2E——一个大规模且真实的端到端网络安全基准,全面评估AI智能体在漏洞发现、利用代码生成及补丁生成全生命周期中的能力。该基准具备全面性和可扩展性,我们构建了自动化、智能体增强的流水线,将开源漏洞数据转化为真实评估环境。目前,基准包含来自139个不同开源项目的920个真实漏洞。
原文摘要 · Abstract (English)
AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabilities. However, existing cybersecurity evaluations of AI systems are limited in scale or scope, and fail to capture the end-to-end lifecycle of real-world software vulnerability discovery and remediation. To address this gap, we propose CyberGym-E2E, a large-scale and realistic end-to-end cybersecurity benchmark that comprehensively evaluates AI agents' abilities across the full lifecycle of vulnerability discovery, PoC generation, and patch generation. CyberGym-E2E is comprehensive and scalable, as we build an automated, agent-enhanced pipeline for transforming open-source vulnerability data into realistic evaluation environments. Currently, the benchmark consists of 920 real-world vulnerabilities across 139 different open-source projects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。