对比多种智能体架构,发现协同方式影响安全检测效果与成本
Towards Optimal Agentic Architectures for Offensive Security Tasks

- 设计20个交互式靶标,测试五类架构在不同模式下的表现
- 白盒模式下验证检测率最高达67.0%,但跨域效果差异显著
- 非单调成本-质量曲线表明:协同并非越广越好,需权衡效率与开销
智能体安全系统越来越多地使用工具型大模型审计实时目标,但以往系统固定单一协调拓扑,难以判断何时增加智能体能提升性能,何时仅增加成本。本文将拓扑选择视为实证系统问题。引入一个包含20个交互式靶标的受控基准(10个网络/API和10个二进制),每个靶标暴露一个可访问端点的真漏洞,分别在白盒和黑盒模式下评估。核心研究共执行600次实验,覆盖五种架构家族、三种模型家族及两种访问模式,另附60次长上下文预研实验。在完整基准上,检测任意率达到58.0%,验证检测率为49.8%。MAS-Indep取得最高验证检测率(64.2%),SAS则为最高效基线,每项有效发现成本仅$0.058。白盒模式显著优于黑盒(67.0% vs. 32.7%),网络任务优于二进制任务(74.3% vs. 25.3%)。置信区间与目标级差异分析表明,可观测性和领域是主要影响因素,部分领先白盒架构统计上无显著差异。主要结论是非单调的成本-质量前沿:更广的协同可提升覆盖率,但在考虑延迟、令牌成本和漏洞验证难度后,并不占优。
原文摘要 · Abstract (English)
Agentic security systems increasingly audit live targets with tool-using LLMs, but prior systems fix a single coordination topology, leaving unclear when additional agents help and when they only add cost. We treat topology choice as an empirical systems question. We introduce a controlled benchmark of 20 interactive targets (10 web/API and 10 binary), each exposing one endpoint-reachable ground-truth vulnerability, evaluated in whitebox and blackbox modes. The core study executes 600 runs over five architecture families, three model families, and both access modes, with a separate 60-run long-context pilot reported only in the appendix. On the completed core benchmark, detection-any reaches 58.0% and validated detection reaches 49.8%. MAS-Indep attains the highest validated detection rate (64.2%), while SAS is the strongest efficiency baseline at $0.058 per validated finding. Whitebox materially outperforms blackbox (67.0% vs. 32.7% validated detection), and web materially outperforms binary (74.3% vs. 25.3%). Bootstrap confidence intervals and paired target-level deltas show that the dominant effects are observability and domain, while some leading whitebox topologies remain statistically close. The main result is a non-monotonic cost-quality frontier: broader coordination can improve coverage, but it does not dominate once latency, token cost, and exploit-validation difficulty are taken into account.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。