用自主LLM代理自动完成渗透测试,提升效率与频率。
AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents
- 基于GPT-4o和LangChain构建自主渗透测试代理
- 在3个HTB靶机上完成15%-25%子任务,略优于人工使用
- 适合安全团队探索自动化漏洞管理新路径
近期研究关注将大语言模型(LLMs)应用于渗透测试,有望降低费用并提高测试频率。本文综述相关工作,识别最佳实践与常见评估问题,并提出AutoPentest——一个基于OpenAI的GPT-4o与LangChain框架的高自主性黑盒渗透测试系统。该系统可执行复杂多步任务,结合外部工具与知识库。我们在三个攻防类Hack The Box(HTB)机器上进行实验,对比了AutoPentest与手动使用ChatGPT-4o界面的基线方法。两者均能完成15%-25%的子任务,其中AutoPentest表现略优。全部实验总成本为96.20美元,而一个月ChatGPT Plus订阅费为20美元。结果表明,未来通过进一步优化及采用更强力的LLM,该方案有望成为漏洞管理的有效组成部分。
原文摘要 · Abstract (English)
A recent area of increasing research is the use of Large Language Models (LLMs) in penetration testing, which promises to reduce costs and thus allow for higher frequency. We conduct a review of related work, identifying best practices and common evaluation issues. We then present AutoPentest, an application for performing black-box penetration tests with a high degree of autonomy. AutoPentest is based on the LLM GPT-4o from OpenAI and the LLM agent framework LangChain. It can perform complex multi-step tasks, augmented by external tools and knowledge bases. We conduct a study on three capture-the-flag style Hack The Box (HTB) machines, comparing our implementation AutoPentest with the baseline approach of manually using the ChatGPT-4o user interface. Both approaches are able to complete 15-25 % of the subtasks on the HTB machines, with AutoPentest slightly outperforming ChatGPT. We measure a total cost of \$96.20 US when using AutoPentest across all experiments, while a one-month subscription to ChatGPT Plus costs \$20. The results show that further implementation efforts and the use of more powerful LLMs released in the future are likely to make this a viable part of vulnerability management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。