arXiv:2504.10112cs.CRcs.AI2025-04被引 20

评估大模型攻防测试的基准方法,指出当前研究缺陷并提改进建议。

Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

  • 分析19篇论文的测试环境与评估指标设计
  • 发现多数研究缺乏真实场景模拟和统一基线
  • 适合安全研究者与攻防工具开发者参考

大语言模型(LLMs)已成为驱动进攻性渗透测试工具的强大手段。由于LLM的黑箱特性,其有效性通常通过实证方法评估,而评估质量高度依赖所选测试环境、度量指标及分析方法。本文系统分析了用于评估大语言模型驱动攻击的方法论与基准实践,聚焦于网络安全中的进攻性应用。我们回顾了18个原型及其对应测试环境的19篇研究论文,总结发现,并为未来研究提供可操作建议,强调扩展现有测试环境、建立基准线、引入全面的度量指标与定性分析的重要性。同时指出安全研究与实际应用之间的差异,提示基于CTF的挑战可能无法完全反映真实渗透测试场景。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as a powerful approach for driving offensive penetration-testing tooling. Due to the opaque nature of LLMs, empirical methods are typically used to analyze their efficacy. The quality of this analysis is highly dependent on the chosen testbed, captured metrics and analysis methods employed. This paper analyzes the methodology and benchmarking practices used for evaluating Large Language Model (LLM)-driven attacks, focusing on offensive uses of LLMs in cybersecurity. We review 19 research papers detailing 18 prototypes and their respective testbeds. We detail our findings and provide actionable recommendations for future research, emphasizing the importance of extending existing testbeds, creating baselines, and including comprehensive metrics and qualitative analysis. We also note the distinction between security research and practice, suggesting that CTF-based challenges may not fully represent real-world penetration testing scenarios.

大模型安全攻防测试基准评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。