首个评估大模型在变电站环境网络安全能力的框架,填补了工业控制领域评测空白。
CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments
- 构建针对IEC 61850协议的专用工具栈,支持工业场景交互
- 5个主流大模型在静态分析任务中表现良好,动态任务性能显著下降
- 适合工业安全研究者、智能电网开发者参考,推动AI在OT安全中的可信应用
大型语言模型(LLM)的快速发展引发了其在网络安全中双重用途的担忧。现有评估框架主要聚焦于信息技术(IT)环境,未能涵盖运营技术(OT)的约束条件和专用协议。为此,我们提出CritBench,一个专用于评估大模型代理在IEC 61850数字变电站环境中网络安全能力的新框架。我们在81项领域特定任务上评估了五个先进模型,包括OpenAI的GPT-5系列与开源模型,涵盖静态配置分析、网络流量侦察及实时虚拟机交互。为支持工业协议交互,我们开发了专用工具栈。实验表明,代理在静态结构化文件分析和单工具网络枚举任务中表现可靠,但在动态任务中性能下降。尽管模型具备对IEC 61850标准术语的显式内化知识,仍难以完成持续的序列推理与状态追踪以操控实时系统,缺乏专用工具时尤为明显。引入我们的领域工具栈可显著缓解此操作瓶颈。代码与评估脚本已公开:https://github.com/GKeppler/CritBench
原文摘要 · Abstract (English)
The advancement of Large Language Models (LLMs) has raised concerns regarding their dual-use potential in cybersecurity. Existing evaluation frameworks overwhelmingly focus on Information Technology (IT) environments, failing to capture the constraints, and specialized protocols of Operational Technology (OT). To address this gap, we introduce CritBench, a novel framework designed to evaluate the cybersecurity capabilities of LLM agents within IEC 61850 Digital Substation environments. We assess five state-of-the-art models, including OpenAI's GPT-5 suite and open-weight models, across a corpus of 81 domain-specific tasks spanning static configuration analysis, network traffic reconnaissance, and live virtual machine interaction. To facilitate industrial protocol interaction, we develop a domain-specific tool scaffold. Our empirical results show that agents reliably execute static structured-file analysis and single-tool network enumeration, but their performance degrades on dynamic tasks. Despite demonstrating explicit, internalized knowledge of the IEC 61850 standards terminology, current models struggle with the persistent sequential reasoning and state tracking required to manipulate live systems without specialized tools. Equipping agents with our domain-specific tool scaffold significantly mitigates this operational bottleneck. Code and evaluation scripts are available at: https://github.com/GKeppler/CritBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。