arXiv:2411.08813cs.AI2024-11被引 4

用大模型分析网络安全评估的缺陷,发现并改进现有评测方法

Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique

  • 用大模型系统性检视Meta网络安全评测方法的漏洞
  • 指出其检测不安全代码时存在显著误判与漏判问题
  • 为后续评测体系优化提供可复现的智能分析范式

网络安全评估领域的重要进展来自Meta提出的CyberSecEval方法。尽管这项工作对新兴领域贡献显著,但其方法仍存在明显局限,尤其体现在不安全代码检测环节。本文深入剖析这些缺陷,并以此为案例,探索大语言模型在基准测试分析中的辅助作用,验证其在识别评测漏洞、提升评估可靠性方面的潜力。

原文摘要 · Abstract (English)

A key development in the cybersecurity evaluations space is the work carried out by Meta, through their CyberSecEval approach. While this work is undoubtedly a useful contribution to a nascent field, there are notable features that limit its utility. Key drawbacks focus on the insecure code detection part of Meta's methodology. We explore these limitations, and use our exploration as a test case for LLM-assisted benchmark analysis.

网络安全大模型评测分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。