arXiv:2605.10834cs.AIcs.CR2026-05

新评测协议让AI渗透测试更贴近真实世界,发现漏洞才是硬道理。

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

论文配图:From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
图 1 · 摘自论文原文
  • 用语义匹配和双部匹配法,精准识别真实目标中的漏洞
  • 在多攻击面、多漏洞类型的复杂环境中验证AI表现
  • 支持持续评估与高效实验,适合安全研究者实战测试

AI渗透测试代理作为进攻性安全系统日益可信,但现有基准对真实场景下哪个表现最佳仍缺乏指导。现有评估协议聚焦于预设目标,如夺旗、远程代码执行、漏洞复现或轨迹相似性,在简化或狭窄环境中进行。这些工具虽能衡量有限能力,却无法充分反映真实渗透中所需的复杂性、开放探索与策略决策。本文提出一种实用的评估协议,将评估重点从任务完成转向经验证的漏洞发现,可在涵盖多个攻击面和漏洞类别的复杂目标上实施。该协议结合结构化真值与基于大模型的语义匹配识别漏洞,采用双部匹配评分机制应对现实中的模糊性,支持持续更新真值、对随机性代理进行重复累积评估、效率度量,并通过精简套件实现可持续实验。该协议显著提升了对AI渗透测试代理的现实可比性与操作信息量。为确保可复现性,我们还发布了专家标注的真值数据与评估协议代码:https://github.com/ethiack/ethibench。

原文摘要 · Abstract (English)

AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for predefined goals such as capture-the-flag, remote code execution, exploit reproduction, or trajectory similarity, in simplified or narrow settings. These tools are valuable for measuring bounded capabilities, yet they do not adequately capture the complexity, open-ended exploration, and strategic decision-making required in realistic pentesting. In this paper, we present a practical evaluation protocol that shifts assessment from task completion to validated vulnerability discovery, allowing evaluation in sufficiently complex targets spanning multiple attack surfaces and vulnerability classes. The protocol combines structured ground-truth with LLM-based semantic matching to identify vulnerabilities, bipartite resolution to score findings under realistic ambiguity, continuous ground-truth maintenance, repeated and cumulative evaluation of stochastic agents, efficiency metrics, and reduced-suite selection for sustainable experimentation. This protocol extends the state of the art by enabling a more realistic, operationally informative comparison of AI pentesting agents. To enable reproducibility, we also release expert-annotated ground truth and code for the proposed evaluation protocol: https://github.com/ethiack/ethibench.

AI安全渗透测试漏洞发现评估协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。