arXiv:2604.17159cs.CRcs.AI2026-04被引 2

对比10个顶级大模型在攻防网络安全任务中的表现,发现环境工具比提示工程更重要。

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks

论文配图:Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
图 1 · 摘自论文原文
  • 构建包含100+渗透工具的Kali Linux环境,支持多模型协作评测
  • Claude 4.5 Opus解决率最高达59%,Gemini 3 Flash成本最低仅0.05美元/解题
  • 提示优化在装备完善时反而降低效果,同模型搭配优于混合配置

我们首次对前沿大语言模型在进攻性网络安全任务上的能力进行了最全面的跨模型评估,基于纽约大学CTF基准测试(NYU CTF Bench)的全部200个挑战,评测了来自7家厂商的10个先进模型。依托D-CIPHER多智能体框架,扩展支持多供应商后端、定制化包含100余种渗透测试工具的Kali Linux环境及运行时工具发现代理。通过受控因子实验发现:相较于Ubuntu环境,使用Kali Linux带来9.5个百分点的性能提升;而自动提示与分类提示常在资源充足的环境下降低表现。模型中,Claude 4.5 Opus解决率最高(59%),紧随其后的是Gemini 3 Pro(52%),Gemini 3 Flash则以每解题0.05美元的成本表现出最优性价比。异构规划者/执行者模型组合未带来显著优势,一致使用相同模型的配置始终优于不同层级模型配对。结果表明,环境工具配置和模型选择是性能的最强驱动因素,而提示工程在条件优越时反而呈现边际递减甚至负收益。报告性能同时反映模型推理能力及其与代理工具链和API集成的兼容性。

原文摘要 · Abstract (English)

We present, to our knowledge, the most comprehensive cross-model evaluation of LLM agents on offensive cybersecurity tasks, benchmarking 10 frontier models from 7 providers on all 200 challenges of the NYU CTF Bench. Building on the D-CIPHER multi-agent framework, we extend it with multi-provider backend support, a custom Kali Linux environment with over 100 pre-installed penetration testing tools, and runtime tool-discovery agents. Through a controlled factorial study, we find that the Kali Linux environment yields a +9.5 percentage-point improvement over Ubuntu, while auto-prompting and category-specific tips often degrade performance in well-equipped environments. Among models, Claude 4.5 Opus achieves the highest solve rate (59%), followed by Gemini 3 Pro (52%), with Gemini 3 Flash offering the best cost-efficiency at $0.05 per solve. Asymmetric planner/executor model assignments provide no meaningful benefit while coherent same-model configurations consistently outperform mixed-tier pairings. Our results indicate that environment tooling and model selection emerge as the strongest drivers of performance, whereas prompt engineering interventions show diminishing or negative returns in well-equipped environments. Reported performance reflects both model reasoning ability and compatibility with agent tooling and API integration.

大模型评测安全攻防工具链集成成本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。