arXiv:2606.14295cs.CRcs.AI2026-06被引 3

首个开放的多主机攻防测试平台,评估前沿AI在真实网络环境中的攻击能力。

AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges

论文配图:AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges
图 1 · 摘自论文原文
  • 构建包含15个真实应用、156台内网主机的多场景攻防环境,模拟真实入侵流程。
  • GPT-5.5+Codex在真实攻防中完成31.7%的后续渗透任务,提示增强后提升至46.3%。
  • 发现未公开漏洞与绕过防御的载荷变异,揭示潜在新型威胁风险。

前沿AI系统在代码审查、漏洞检测和利用方面能力日益增强,但其进攻能力评估仍受限于缺乏开放、可复现的多主机攻防环境。现有公开基准多聚焦孤立技能如CTF求解、漏洞复现和利用生成,却忽略了真实入侵流程:发现暴露服务、获取初始立足点、收集内部信息并横向扩展。这一差距导致难以早期发现新兴风险。本文提出AgentCyberRange,首个开源的多范围基础设施,用于衡量智能体在真实攻防环境中自主攻击的能力。它整合了110个漏洞、15个真实Web应用及8个企业级攻防环境,共156台内网主机,并配备Cage工具链实现执行、编排、结果收集与验证。基准涵盖两大核心阶段:网页攻防(探索暴露应用并验证漏洞)与事后渗透(从初始立足点扩大内网控制)。我们在统一提示与预算下评估六款前沿AI系统。GPT-5.5搭配Codex表现最佳,解决16.1%的网页攻防任务与31.7%的事后渗透任务;提供更具体提示后,该比例分别提升至33.0%和46.3%。还观察到未披露漏洞及可绕过主机防御的载荷变异等非基准发现。结果表明,在真实、可复现条件下开展开放攻防评估,对洞察前沿AI的潜在进攻能力至关重要。

原文摘要 · Abstract (English)

Frontier AI systems are increasingly capable of cybersecurity tasks, including codebase inspection, vulnerability detection, and exploitation. However, evaluating their offensive capabilities remains constrained by limited access to open, reproducible, multi-host cyber ranges. Existing public benchmarks capture isolated skills such as CTF solving, vulnerability reproduction, and exploit generation, but often abstract away realistic intrusion workflows: discovering exposed services, gaining a foothold, collecting internal information, and expanding compromise across hosts. This gap makes it difficult to observe emerging risks early, because frontier AI systems are rarely evaluated under realistic attack conditions. We introduce AgentCyberRange, the first open, multi-range infrastructure for measuring autonomous cyber attack capability in realistic cyber ranges. It combines 110 vulnerabilities across 15 real web applications and 8 enterprise-like cyber ranges with 156 internal hosts, plus Cage, a toolchain for execution, orchestration, result collection, and verification. The benchmark covers two core stages: web exploitation, where agents explore exposed applications and validate vulnerabilities, and post exploitation, where agents turn an initial foothold into broader internal compromise. We evaluate six frontier AI systems under matched prompts and budgets. GPT-5.5 with Codex performs best, solving 16.1% of web exploitation tasks and 31.7% of post-exploitation tasks; with more concrete hints, these rates increase to 33.0% and 46.3%. We also observe out-of-benchmark findings, including unknown vulnerabilities in popular projects, and payload mutation that bypasses host defenses. These results show that open cyber-range evaluation is necessary for observing emerging offensive capabilities under realistic and reproducible conditions.

AI攻防安全评测红队对抗漏洞挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。