arXiv:2509.13021cs.CRcs.AI2025-09被引 6

用大模型驱动的多智能体系统自动攻防测试,效率远超人工。

xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

  • 用微调的大模型分派侦察、扫描、利用等角色协同工作
  • 在两个基准上完成率79.17%,超过VulnBot和PentestGPT
  • 适合安全研究者和自动化测试团队快速部署

本文提出xOffense,一个基于领域适配大语言模型的自主多智能体渗透测试框架,将传统依赖人力、专家经验的手动流程转变为可完全自动化、机器可执行的工作流,能随计算资源弹性扩展。核心采用微调后的中等规模开源大模型Qwen3-32B进行推理与决策,分配专用智能体负责侦察、漏洞扫描与利用,并通过编排层实现各阶段无缝协作。在链式思维渗透测试数据上微调后,模型能生成精确工具指令并保持多步推理一致性。我们在AutoPenBench和AI-Pentest-Benchmark两个严格基准上评估,结果表明xOffense持续优于现有方法,子任务完成率达79.17%,显著超越VulnBot与PentestGPT。研究显示,嵌入结构化多智能体编排中的领域适配中等规模大模型,可提供更优、更低成本、可复现的自主渗透测试解决方案。

原文摘要 · Abstract (English)

This work introduces xOffense, an AI-driven, multi-agent penetration testing framework that shifts the process from labor-intensive, expert-driven manual efforts to fully automated, machine-executable workflows capable of scaling seamlessly with computational infrastructure. At its core, xOffense leverages a fine-tuned, mid-scale open-source LLM (Qwen3-32B) to drive reasoning and decision-making in penetration testing. The framework assigns specialized agents to reconnaissance, vulnerability scanning, and exploitation, with an orchestration layer ensuring seamless coordination across phases. Fine-tuning on Chain-of-Thought penetration testing data further enables the model to generate precise tool commands and perform consistent multi-step reasoning. We evaluate xOffense on two rigorous benchmarks: AutoPenBench and AI-Pentest-Benchmark. The results demonstrate that xOffense consistently outperforms contemporary methods, achieving a sub-task completion rate of 79.17%, decisively surpassing leading systems such as VulnBot and PentestGPT. These findings highlight the potential of domain-adapted mid-scale LLMs, when embedded within structured multi-agent orchestration, to deliver superior, cost-efficient, and reproducible solutions for autonomous penetration testing.

渗透测试多智能体大模型应用自动化安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。