用大模型自省式渗透测试,比普通模型多发现16.7%凭证
RefPentester: A Knowledge-Informed Self-Reflective Penetration Testing Framework Based on Large Language Models
- 基于七阶段状态机,让大模型自我反思失败经验
- 在Hack The Box的Sau机器上成功提取凭证,胜过GPT-4o 16.7%
- 适合安全人员提升自动化渗透效率
由大语言模型驱动的自动化渗透测试(AutoPT)因能利用模型内嵌知识自动识别系统漏洞而受到关注。然而,现有基于LLM的AutoPT框架在复杂任务中表现仍不及人类专家,原因包括训练知识分布不均、规划过程短视以及命令生成时的幻觉。此外,渗透测试的试错特性受限于缺乏从失败中学习的机制,难以实现策略迭代优化。为此,我们提出一种基于大模型的知识引导、自省式渗透测试框架RefPentester。该框架可协助人类操作员判断当前渗透阶段,选择合适战术与技术,推荐操作步骤,并对过往失败进行反思与学习。我们还构建了七状态阶段机模型以有效整合该框架。评估显示,RefPentester在Hack The Box的Sau机器上成功暴露凭证,较基线GPT-4o模型提升16.7%;在各阶段过渡成功率上也表现更优。
原文摘要 · Abstract (English)
Automated penetration testing (AutoPT) powered by large language models (LLMs) has gained attention for its ability to automate ethical hacking processes and identify vulnerabilities in target systems by leveraging the inherent knowledge of LLMs. However, existing LLM-based AutoPT frameworks often underperform compared to human experts in challenging tasks for several reasons: the imbalanced knowledge used in LLM training, short-sightedness in the planning process, and hallucinations during command generation. Moreover, the trial-and-error nature of the PT process is constrained by existing frameworks lacking mechanisms to learn from previous failures, restricting adaptive improvement of PT strategies. To address these limitations, we propose a knowledge-informed, self-reflective PT framework powered by LLMs, called RefPentester. This AutoPT framework is designed to assist human operators in identifying the current stage of the PT process, selecting appropriate tactics and techniques for each stage, choosing suggested actions, providing step-by-step operational guidance, and reflecting on and learning from previous failed operations. We also modeled the PT process as a seven-state Stage Machine to integrate the proposed framework effectively. The evaluation shows that RefPentester can successfully reveal credentials on Hack The Box's Sau machine, outperforming the baseline GPT-4o model by 16.7%. Across PT stages, RefPentester also demonstrates superior success rates on PT stage transitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。