arXiv:2605.04499cs.CRcs.AI2026-05被引 1

用逻辑推理生成渗透测试策略,提升自动化安全检测效果。

Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis

论文配图:Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis
图 1 · 摘自论文原文
  • 构建领域专用推理模型,通过逻辑推导生成渗透策略。
  • 策略生成性能提升87%,子任务完成率提高47.5%。
  • 适合安全研究者与自动化渗透工具开发者使用。

网络威胁快速增加,影响范围从大型企业扩展至政府服务与个人用户,亟需强大安全系统。然而,熟练的网络安全人才严重短缺,加剧了挑战。尽管已有研究尝试用大模型代理实现渗透测试自动化,但现有框架因策略制定能力弱、领域推理不足及动作与工具选择不准而表现不佳。为此,我们提出Pen-Strategist框架,包含一个基于逻辑推理生成渗透策略的新颖领域特定模型,以及一个将策略转化为可执行步骤的分类器。首先,构建包含渗透测试场景中策略推导与步骤选择逻辑解释的推理数据集;随后,使用强化学习微调Qwen-3-14B模型以生成策略。在数据集测试集上的评估显示,策略生成性能相比基线提升87%。进一步将微调后的Pen-Strategist集成至PentestGPT等自动化渗透框架,在漏洞机器上测试,子任务完成率提升47.5%,超越基线GPT-5。在CTFKnow基准测试中,性能比基线模型高18%。对于步骤预测,训练基于语义的CNN分类器,其表现优于商用大模型28%,并提升执行稳定性。最后,通过用户研究定性评估生成策略,Pen-Strategist在表现上优于Claude-4.6-Sonnet。

原文摘要 · Abstract (English)

Cyber threats are rapidly increasing, expanding their impact from large-scale enterprises to government services and individual users, making robust security systems increasingly essential. However, a significant shortage of skilled cybersecurity professionals exacerbates this challenge. While recent research has explored automating tasks such as penetration testing using LLM-based agents, existing frameworks often perform poorly due to limited capability in strategy formulation, domain-specific reasoning, and accurate action and tool selection. To overcome these limitations, we propose Pen-Strategist framework, consisting of a novel domain-specific reasoning model that derives pentesting strategies via logical reasoning and a classifier that converts the strategies into actionable steps. First, we construct a reasoning dataset containing logical explanations for both strategy derivation and step selection in pentesting scenarios. We then fine-tune a Qwen-3-14B model for strategy generation using reinforcement learning. Evaluation on the test split of the dataset demonstrates a 87% improvement in strategy derivation performance compared to the baseline. Furthermore, we integrate the fine-tuned Pen-Strategist model into existing automated pentesting frameworks, such as PentestGPT, and evaluate its performance on vulnerable machines, achieving a 47.5% improvement in subtask completion while surpassing the baseline GPT-5. Further experiments on the CTFKnow benchmark show an 18% performance gain over the base model. For step prediction, we train a semantic-based CNN classifier, which outperforms commercial LLMs by 28% and enhances execution stability. Finally, we conduct a user study to qualitatively assess the generated strategies, and Pen-Strategist demonstrates superior performance compared to the Claude-4.6-Sonnet.

渗透测试大模型安全分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。