首个分阶段评测大模型渗透测试能力的基准,揭示现有系统成功率仅31%。
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
- 按信息收集、漏洞筛选、攻击决策等6个阶段拆解渗透流程
- 9个主流大模型在346个任务中平均成功率仅31%,自主代理几乎失效
- 提供可复现的自动化评估框架,适合安全与AI交叉研究者
渗透测试对评估和增强系统安全性至关重要,但传统流程高度依赖人工、需专业知识且难以扩展。尽管大语言模型(LLMs)为自动化带来希望,现有应用多采用简单提示,缺乏任务分解与领域适配,导致行为不可靠且无法深入理解模型在各阶段的能力。为此,我们提出PentestEval,首个针对渗透测试六阶段(信息收集、弱点获取与过滤、攻击决策、漏洞利用生成与修改)的综合性基准。该基准整合专家标注的真值数据,通过全自动评估流水线覆盖12个真实漏洞场景中的346个任务。对9种广泛使用的LLM进行分阶段评估发现,整体表现薄弱且各阶段局限明显;端到端流程成功率仅为31%,现有系统如PentestGPT、PentestAgent和VulnBot也存在类似问题,自主代理几乎完全失败。结果表明,自主渗透测试需要更强的结构化推理能力,模块化设计能提升各阶段性能并改善整体效果。PentestEval为未来细粒度、阶段级评估奠定了基础,推动更可靠的基于LLM的自动化发展。
原文摘要 · Abstract (English)
Penetration testing is essential for assessing and strengthening system security against real-world threats, yet traditional workflows remain highly manual, expertise-intensive, and difficult to scale. Although recent advances in Large Language Models (LLMs) offer promising opportunities for automation, existing applications rely on simplistic prompting without task decomposition or domain adaptation, resulting in unreliable black-box behavior and limited insight into model capabilities across penetration testing stages. To address this gap, we introduce PentestEval, the first comprehensive benchmark for evaluating LLMs across six decomposed penetration testing stages: Information Collection, Weakness Gathering and Filtering, Attack Decision-Making, Exploit Generation and Revision. PentestEval integrates expert-annotated ground truth with a fully automated evaluation pipeline across 346 tasks covering all stages in 12 realistic vulnerable scenarios. Our stage-level evaluation of 9 widely used LLMs reveals generally weak performance and distinct limitations across the stages of penetration-testing workflow. End-to-end pipelines reach only 31% success rate, and existing LLM-powered systems such as PentestGPT, PentestAgent, and VulnBot exhibit similar limitations, with autonomous agents failing almost entirely. These findings highlight that autonomous penetration testing demands stronger structured reasoning, where modularization enhances each individual stage and improves overall performance. PentestEval provides the foundational benchmark needed for future research on fine-grained, stage-level evaluation, paving the way toward more reliable LLM-based automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。