arXiv:2604.05719cs.CRcs.AI2026-04被引 3

首次系统分析大模型渗透测试框架,揭示其设计与实际表现

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

  • 从六维度梳理现有框架架构设计
  • 13个开源框架实测生成超1500条日志
  • 适合安全研究者和自动化测试开发者

大型语言模型的快速发展为自动化渗透测试(AutoPT)带来了新机遇,催生了众多旨在实现端到端自主攻击的框架。然而,现有研究普遍缺乏系统的架构分析和统一基准下的大规模实证比较。本文首次提出面向该领域的知识体系化(SoK),从代理架构、规划、记忆、执行、外部知识和基准六个维度全面回顾现有框架设计。在实证层面,我们在统一基准上对13个代表性开源AutoPT框架及2个基线框架进行了大规模实验,总计消耗超过100亿个标记,生成超过1500条执行日志,并由15位以上具备网络安全专长的研究人员历时四个月进行人工审查与分析。通过深入调研该快速发展的领域最新进展,我们为研究者提供了结构化的分类体系、大规模实证基准及未来研究方向。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has created new opportunities for Automated Penetration Testing (AutoPT), spawning numerous frameworks aimed at achieving end-to-end autonomous attacks. However, despite the proliferation of related studies, existing research generally lacks systematic architectural analysis and large-scale empirical comparisons under a unified benchmark. Therefore, this paper presents the first Systematization of Knowledge (SoK) focusing on the architectural design and comprehensive empirical evaluation of current LLM-based AutoPT frameworks. At systematization level, we comprehensively review existing framework designs across six dimensions: agent architecture, agent plan, agent memory, agent execution, external knowledge, and benchmarks. At empirical level, we conduct large-scale experiments on 13 representative open-source AutoPT frameworks and 2 baseline frameworks utilizing a unified benchmark. The experiments consumed over 10 billion tokens in total and generated more than 1,500 execution logs, which were manually reviewed and analyzed over four months by a panel of more than 15 researchers with expertise in cybersecurity. By investigating the latest progress in this rapidly developing field, we provide researchers with a structured taxonomy to understand existing LLM-based AutoPT frameworks and a large-scale empirical benchmark, along with promising directions for future research.

自动化测试大模型安全渗透测试LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。