对比三款AI工具在渗透测试中的表现,发现其能显著提升效率但无法完全替代人工。
Generative Artificial Intelligence-Supported Pentesting: A Comparison between Claude Opus, GPT-4, and Copilot
- 在标准渗透测试流程中测试三款主流AI工具
- Claude Opus在各阶段表现最优,尤其在漏洞挖掘与报告生成上领先
- 适合安全研究人员快速辅助测试,但需人工验证结果
生成式人工智能(GenAI)正深刻影响社会,尤其在网络安全领域具有重要应用价值。本文聚焦于其在渗透测试(pentesting)或道德黑客攻击中的应用潜力,评估了当前主流通用型GenAI工具——Claude Opus、GPT-4(ChatGPT)和Copilot——在符合渗透测试执行标准(PTES)的全流程中的表现。实验在受控虚拟环境中进行,涵盖所有标准阶段。结果显示,尽管这些工具尚无法完全自动化渗透测试,但在特定任务中显著提升了效率与效果。三者均具实用性,其中Claude Opus在多数场景下表现最佳,尤其在漏洞识别与测试报告生成方面优势明显。
原文摘要 · Abstract (English)
The advent of Generative Artificial Intelligence (GenAI) has brought a significant change to our society. GenAI can be applied across numerous fields, with particular relevance in cybersecurity. Among the various areas of application, its use in penetration testing (pentesting) or ethical hacking processes is of special interest. In this paper, we have analyzed the potential of leading generic-purpose GenAI tools-Claude Opus, GPT-4 from ChatGPT, and Copilot-in augmenting the penetration testing process as defined by the Penetration Testing Execution Standard (PTES). Our analysis involved evaluating each tool across all PTES phases within a controlled virtualized environment. The findings reveal that, while these tools cannot fully automate the pentesting process, they provide substantial support by enhancing efficiency and effectiveness in specific tasks. Notably, all tools demonstrated utility; however, Claude Opus consistently outperformed the others in our experimental scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。