arXiv:2508.01054cs.CRcs.AI2025-08被引 5

GPT-4o可自主完成80%的初阶渗透测试任务,效率超人类。

Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs

  • 将GPT-4o接入SSH框架,直接生成命令解决安全挑战
  • 25个关卡中成功18个,加提示再解2个,总成功率80%
  • 适合想快速入门渗透测试或提升效率的安全研究者

本研究评估GPT-4o在连接OverTheWire Bandit攻防竞赛环境时,自主解决初级攻击任务的能力。在25个支持单命令SSH框架的关卡中,GPT-4o无需辅助成功解决18个,另两个经极简提示后完成,整体成功率80%。模型在涉及Linux文件系统导航、数据提取或解码、基础网络操作的单步挑战中表现优异,常一击命中正确命令,速度超过人类。失败案例多出现在需多命令协作、持久工作目录、复杂网络侦察、守护进程创建或非标准Shell交互等场景。这些局限反映当前架构缺陷,而非缺乏漏洞知识。结果表明大语言模型(LLMs)可自动化大量新手渗透测试流程,降低攻击门槛,同时为防御方提供高效侦察辅助。未解任务揭示了安全设计环境中可能阻碍简单LLM攻击的薄弱点,有助于未来加固策略制定。此外,成果还暗示将LLMs用于网络安全教育实践的潜力。

原文摘要 · Abstract (English)

This study evaluates the ability of GPT-4o to autonomously solve beginner-level offensive security tasks by connecting the model to OverTheWire's Bandit capture-the-flag game. Of the 25 levels that were technically compatible with a single-command SSH framework, GPT-4o solved 18 unaided and another two after minimal prompt hints for an overall 80% success rate. The model excelled at single-step challenges that involved Linux filesystem navigation, data extraction or decoding, and straightforward networking. The approach often produced the correct command in one shot and at a human-surpassing speed. Failures involved multi-command scenarios that required persistent working directories, complex network reconnaissance, daemon creation, or interaction with non-standard shells. These limitations highlight current architectural deficiencies rather than a lack of general exploit knowledge. The results demonstrate that large language models (LLMs) can automate a substantial portion of novice penetration-testing workflow, potentially lowering the expertise barrier for attackers and offering productivity gains for defenders who use LLMs as rapid reconnaissance aides. Further, the unsolved tasks reveal specific areas where secure-by-design environments might frustrate simple LLM-driven attacks, informing future hardening strategies. Beyond offensive cybersecurity applications, results suggest the potential to integrate LLMs into cybersecurity education as practice aids.

渗透测试大模型安全攻防自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。