arXiv:2412.20787cs.CRcs.AI2024-12被引 29

构建首个多维度的网络安全大模型评测数据集,覆盖中英文、多种题型和能力层级。

SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity

  • 设计多维度评测框架,包含选择题与简答题,覆盖知识记忆与逻辑推理。
  • 收集44,823道选择题和3,087道简答题,涵盖多个网络安全子领域。
  • 用大模型自动标注与评分,兼顾质量与效率,适合安全领域研究者使用。

评估大语言模型(LLMs)对理解其在自然语言处理和代码生成等应用中的能力与局限至关重要。现有基准如MMLU、C-Eval和HumanEval虽能评估通用性能,但缺乏对网络安全等专业领域的关注。此前的网络安全数据集存在数据量不足、依赖单选题等问题。为此,我们提出SecBench,一个用于评估大模型在网络安全领域表现的多维度基准数据集。该数据集包含多种题型(单选题与简答题)、不同能力层级(知识保留与逻辑推理)、多语言支持(中文与英文)及多个子领域。数据通过开源渠道收集,并举办网络安全题目设计大赛获得,共含44,823道单选题和3,087道简答题。特别地,我们利用高效的大模型完成数据标注与构建自动评分代理,实现对简答题的自动化评估。在16个主流大模型上的测试结果表明,SecBench具有良好的可用性,是目前规模最大、最全面的网络安全领域大模型评测数据集。

原文摘要 · Abstract (English)

Evaluating Large Language Models (LLMs) is crucial for understanding their capabilities and limitations across various applications, including natural language processing and code generation. Existing benchmarks like MMLU, C-Eval, and HumanEval assess general LLM performance but lack focus on specific expert domains such as cybersecurity. Previous attempts to create cybersecurity datasets have faced limitations, including insufficient data volume and a reliance on multiple-choice questions (MCQs). To address these gaps, we propose SecBench, a multi-dimensional benchmarking dataset designed to evaluate LLMs in the cybersecurity domain. SecBench includes questions in various formats (MCQs and short-answer questions (SAQs)), at different capability levels (Knowledge Retention and Logical Reasoning), in multiple languages (Chinese and English), and across various sub-domains. The dataset was constructed by collecting high-quality data from open sources and organizing a Cybersecurity Question Design Contest, resulting in 44,823 MCQs and 3,087 SAQs. Particularly, we used the powerful while cost-effective LLMs to (1). label the data and (2). constructing a grading agent for automatic evaluation of SAQs. Benchmarking results on 16 SOTA LLMs demonstrate the usability of SecBench, which is arguably the largest and most comprehensive benchmark dataset for LLMs in cybersecurity. More information about SecBench can be found at our website, and the dataset can be accessed via the artifact link.

大模型评测网络安全多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。