arXiv:2604.20389cs.CRcs.AI2026-04

评测大模型在网络安全认证知识上的表现,发现其在通用领域接近专家水平。

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge

  • 构建基于行业认证的多选题评测集,覆盖IT与工控安全等领域。
  • 顶尖模型在通用网络安全知识上达人类专家水平,但对特定厂商或标准细节准确率下降。
  • 提出可解释的生成框架,为模型决策提供自然语言理由,适合安全领域研究者使用。

大型语言模型在专业工作流中的快速应用,亟需对其领域知识进行评估。本文提出CyberCertBench,一个源自行业认证的多选题问答评测集,涵盖信息技术网络安全及工业控制系统等细分领域。该评测对比模型在信息安全与相关标准(如IEC 62443)方面的表现。同时,我们提出并验证了一种新的提议-验证框架,用于生成可解释的自然语言推理过程。评估显示,前沿模型在通用网络与信息安全知识上已达到人类专家水平;但在涉及厂商特定细节或正式标准的问题上,准确率显著下降。模型规模与发布时间分析表明,参数效率持续提升,而近期大模型出现收益递减现象。代码与评测脚本已在GitHub公开。

原文摘要 · Abstract (English)

The rapid evolution and use of Large Language Models (LLMs) in professional workflows require an evaluation of their domain-specific knowledge against industry standards. We introduceCyberCertBench, a new suite of Multiple Choice Question Answering (MCQA) benchmarks derived from industry recognized certifications. CyberCertBench evaluates LLM domain knowledgeagainst the professional standards of Information Technology cybersecurity and more specializedareas such as Operational Technology and related cybersecurity standards. Concurrently, we propose and validate a novel Proposer-Verifier framework, a methodology to generate interpretable,natural language explanations for model performance. Our evaluation shows that frontier modelsachieve human expert level in general networking and IT security knowledge. However, theiraccuracy declines in questions that require vendor-specific nuances or knowledge in formalstandards, like, e.g., IEC 62443. Analysis of model scaling trend and release date demonstratesremarkable gains in parameter efficiency, while recent larger models show diminishing returns.Code and evaluation scripts are available at: https://github.com/GKeppler/CyberCertBench.

大模型评测网络安全可解释性认证基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。