AI安全代理在真实网络渗透测试中表现媲美人类专家,且成本更低。
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
- 构建多智能体框架ARTEMIS,支持动态提示与自动漏洞分类。
- ARTEMIS发现9个有效漏洞,82%提交率,超越9名人类专家。
- 适合关注自动化渗透与成本优化的安全团队参考。
我们首次在真实企业环境中对人工智能代理与人类网络安全专业人员进行综合评估。在包含约8000台主机、分布在12个子网的大型大学网络上,对比了十名网络安全专家与六种现有AI代理及我们新提出的ARTEMIS代理架构。ARTEMIS是一个多智能体框架,具备动态提示生成、任意子代理调用和自动漏洞优先级排序能力。在对比实验中,ARTEMIS总排名第二,共发现9个有效漏洞,有效提交率达82%,优于10名人类参与者中的9人。现有代理如Codex和CyAgent的表现普遍低于多数人类参与者,而ARTEMIS展现出与最强人类相当的技术水平和提交质量。我们观察到AI代理在系统化扫描、并行利用和成本方面具有优势——某些ARTEMIS变体每小时仅需18美元,远低于专业渗透测试人员的60美元。但同时也发现关键短板:AI代理误报率较高,且难以处理基于图形界面的任务。
原文摘要 · Abstract (English)
We present the first comprehensive evaluation of AI agents against human cybersecurity professionals in a live enterprise environment. We evaluate ten cybersecurity professionals alongside six existing AI agents and ARTEMIS, our new agent scaffold, on a large university network consisting of ~8,000 hosts across 12 subnets. ARTEMIS is a multi-agent framework featuring dynamic prompt generation, arbitrary sub-agents, and automatic vulnerability triaging. In our comparative study, ARTEMIS placed second overall, discovering 9 valid vulnerabilities with an 82% valid submission rate and outperforming 9 of 10 human participants. While existing scaffolds such as Codex and CyAgent underperformed relative to most human participants, ARTEMIS demonstrated technical sophistication and submission quality comparable to the strongest participants. We observe that AI agents offer advantages in systematic enumeration, parallel exploitation, and cost -- certain ARTEMIS variants cost $18/hour versus $60/hour for professional penetration testers. We also identify key capability gaps: AI agents exhibit higher false-positive rates and struggle with GUI-based tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。