首个开源自动化渗透测试基准,评估大模型安全能力。
Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
- 构建开放端到端渗透测试基准,评估LLM在真实场景表现。
- GPT-4o与LLama 3.1均无法独立完成全流程渗透测试。
- 揭示枚举、利用、提权等环节的模型瓶颈,指导优化方向。
网络攻击每年造成数十亿美元损失,渗透测试作为防御手段至关重要。尽管大语言模型(LLMs)在多个领域展现潜力,但目前尚无全面、开源、自动化的端到端渗透测试基准来推动研究和评估其安全应用能力。本文提出首个针对基于LLM的自动化渗透测试的开放基准,首次评估GPT-4o与LLama 3.1-405B在先进工具PentestGPT下的表现。结果表明,尽管LLama 3.1略优于GPT-4o,两者仍需人工协助才能完成完整渗透流程。通过消融实验,本文深入分析了模型在枚举、漏洞利用和权限提升等阶段的局限性,为未来改进提供关键洞见。本研究推动了人工智能辅助网络安全的发展,奠定了自动化渗透测试的科研基础。
原文摘要 · Abstract (English)
Hacking poses a significant threat to cybersecurity, inflicting billions of dollars in damages annually. To mitigate these risks, ethical hacking, or penetration testing, is employed to identify vulnerabilities in systems and networks. Recent advancements in large language models (LLMs) have shown potential across various domains, including cybersecurity. However, there is currently no comprehensive, open, automated, end-to-end penetration testing benchmark to drive progress and evaluate the capabilities of these models in security contexts. This paper introduces a novel open benchmark for LLM-based automated penetration testing, addressing this critical gap. We first evaluate the performance of LLMs, including GPT-4o and LLama 3.1-405B, using the state-of-the-art PentestGPT tool. Our findings reveal that while LLama 3.1 demonstrates an edge over GPT-4o, both models currently fall short of performing end-to-end penetration testing even with some minimal human assistance. Next, we advance the state-of-the-art and present ablation studies that provide insights into improving the PentestGPT tool. Our research illuminates the challenges LLMs face in each aspect of Pentesting, e.g. enumeration, exploitation, and privilege escalation. This work contributes to the growing body of knowledge on AI-assisted cybersecurity and lays the foundation for future research in automated penetration testing using large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。