拆分探测与利用,更真实评估大模型的网页渗透能力
Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing
- 将探测和利用分开测试,避免错误传递干扰结果
- 准确提供漏洞信息时,成功率最高达90.0%,自主探测仅50.0%
- 不同架构各有优势,适合不同攻击场景
大型语言模型在自动化渗透测试中展现潜力,但现有端到端黑盒评估易受错误传播影响:早期探测失败会掩盖代理的真实利用能力。为此,我们提出一种两阶段解耦评估框架,将利用执行与探测分离。通过70个高保真网页漏洞测试环境中的真实注入与知识驱动消融实验,该框架将利用性能与探测噪声隔离。我们在50个代表性漏洞的严格对齐子集上评估了五种开源渗透测试代理,涵盖多代理、单体及图驱动架构。结果揭示显著能力差距:在提供准确漏洞上下文时,代理功能成功率达90.0%;而自主探测的针对性漏洞召回率仅约50.0%,主要因解析非结构化遥测失败。跨架构分析进一步发现:多代理设计在长序列交互(如反序列化)中更优,单体与图驱动架构分别在短链注入和跨会话访问控制漏洞上表现更好。该工作提供了细粒度基准协议与下一代自动化攻击代理设计的实证基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown promise for automated penetration testing, yet existing end-to-end black-box evaluations are highly susceptible to error cascading: failures in early reconnaissance can mask an agent's actual ability to exploit vulnerabilities. To more accurately characterize these capabilities, we propose a two-stage decoupled evaluation framework that separates exploit execution from reconnaissance. Using ground-truth injection and knowledge-driven ablation across 70 high-fidelity web vulnerability testbeds, our framework isolates exploitation performance from reconnaissance noise. We empirically evaluate five open-source penetration-testing agents, covering multiagent, monolithic, and graph-driven architectures, on a strictly aligned subset of 50 representative vulnerabilities. The results reveal a substantial capability gap. With accurate vulnerability context, agents achieve a functional success rate of up to 90.0%, whereas autonomous reconnaissance, measured by targeted vulnerability recall, plateaus at approximately 50.0%, primarily due to failures in parsing unstructured telemetry. Cross-architectural analysis further reveals distinct capability niches: multi-agent isolation is more effective for long-sequence interactions such as de-serialization, while monolithic and graph-driven designs perform better on short-chain injections and cross-session access-control vulnerabilities, respectively. This decoupled evaluation work provides a fine-grained benchmarking protocol and an empirical basis for designing next-generation automated offensive security agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。