用大模型自动渗透测试,成功率超84%
APT-Agent: Automated Penetration Testing using Large Language Models

- 用混合修正模块和专用记忆架构解决幻觉与上下文丢失问题
- 在7个漏洞服务上达成84.29%的自动化攻击成功率
- 适合安全研究者和自动化渗透团队快速验证系统弱点
渗透测试对保障现代网络基础设施安全至关重要,但传统人工方法难以应对规模与复杂性。大型语言模型(LLMs)为自动化带来新可能,但现有方法仍存在技术实体幻觉和长期上下文记忆不足两大挑战。为此,我们提出APT-Agent,一个完全自动化的基于LLM的渗透测试框架,系统化地协调侦察、利用和数据外传。APT-Agent引入混合修正模块以恢复幻觉命令,并采用命令级记忆架构,在多步骤攻击序列中保持操作上下文。我们在Metasploitable 2上针对七个跨网页、数据库和网络协议的脆弱服务评估该框架。在相同条件下,APT-Agent实现84.29%的端到端利用成功率,显著高于脚本小子(48.57%)和PentestGPT(18.57%)。通过降低认知负担并减少对人工干预的依赖,APT-Agent代表了可扩展、可靠且认知高效的渗透测试自动化的重要进展。
原文摘要 · Abstract (English)
Penetration testing is essential to securing modern web infrastructures, yet traditional manual methods struggle to keep pace with their scale and complexity. Large Language Models (LLMs) offer new opportunities for automating these tasks, but existing approaches face two persistent challenges: hallucination of technical entities and insufficient long-term contextual memory. To address these issues, we present APT-Agent, a fully automated LLM-driven penetration testing framework that systematically orchestrates reconnaissance, exploitation, and exfiltration. APT-Agent introduces a hybrid rectification module to recover hallucinated commands and a command-specific memory architecture to preserve operational context across multi-step attack sequences. We evaluate our APT-Agent on Metasploitable 2 against seven vulnerable services spanning web, database, and network protocols. APT-Agent achieves an 84.29% end-to-end exploitation success rate, compared to 48.57% (Script Kiddie) and 18.57% (PentestGPT) under matched conditions. By reducing cognitive burden and minimizing reliance on human intervention, APT-Agent represents a step toward scalable, reliable, and cognitively efficient automation for penetration testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。