研究大模型安全工具协同的瓶颈,发现客户端影响远超模型本身。
Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI

- 用150+工具的HexStrikeAI测试86个挑战,对比不同模型与客户端表现
- 解决率从55.4%提升至72.0%,中等难度任务提升最显著
- 失败主因是推理或环境限制,而非缺少工具,适合安全自动化研究者
驱动安全工具集的大语言模型代理正日益普及。然而其能力边界仍不明确:模型与客户端哪个更关键?限定于自身工具是否有益?能力受限于推理还是工具缺失?我们以开源工具协调器HexStrikeAI(支持150+工具)为实验平台,采用评估-诊断-优化的方法,对86个picoCTF挑战(7类、3难度层级)在3种工具访问模式和3种模型/客户端配置下进行测试(共774次试验)。通过修复现有工具、调整代理行为及新增11个功能工具,重新运行此前失败的试验。诊断发现,在固定模型下,客户端是首要影响因素(两个DeepSeek客户端差距达2.1倍),且解题成功率随难度单调下降;中等难度任务提升最大。整体解题率由55.4%升至72.0%,所有配置均显著改善(配对McNemar检验,p < 0.001,95%置信区间无重叠)。剩余失败主要源于推理或环境限制,而非工具缺失。60次稳定性子实验显示单次判决可复现(20次中有17次一致)。我们讨论了评估此类协调器的启示,并明确局限:仅使用单一基准、修复在相同挑战上调优,且客户端效应仅在一种模型上验证,其普适性仍为假设。
原文摘要 · Abstract (English)
Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that bound their capability remain poorly characterized: how much depends on the model versus the client that drives it, whether constraining the agent to the orchestrator's own tools helps, and where capability is limited by reasoning rather than by missing tools. Using HexStrikeAI, an open-source orchestrator that exposes 150+ tools, as a testbed, we follow a methodology that evaluates the system, diagnoses its failures, and applies targeted improvements. We run 86 picoCTF challenges across seven categories and three difficulty tiers, under three tool-access regimes and three model/client configurations (774 trials). We then apply corrections to existing tools, agent-behavior changes, and eleven new capability tools, and re-run the previously-unsuccessful trials. The diagnosis isolates the driving client as a first-order factor for a fixed model (a 2.1 * gap between two DeepSeek clients) and a monotonic difficulty gradient, with the largest gains in the mid tier. The overall solve rate rises from 55.4% to 72.0%, and every configuration improves significantly (paired McNemar p < 0.001, non-overlapping 95% confidence intervals). The residual failures are reasoning- or environment-bound rather than missing-tool. A 60-run stability sub-study finds single-run verdicts reproducible (17/20 unanimous). We discuss what the results imply for how such orchestrators should be evaluated, and we are explicit about the limits: the study uses a single benchmark, the fixes were tuned on the same challenges they were evaluated on, and the client effect is demonstrated for one model only, so its generality to other models remains a hypothesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。