arXiv:2604.23341cs.CRcs.AI2026-04

测试大模型在电网运维中被恶意提示攻破的风险,发现部分模型易遭操控。

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

  • 用三种主流大模型测试三种越狱攻击,在电网标准场景下评估安全性
  • 整体攻击成功率33.1%,其中深度伪装攻击最高达63.17%成功
  • 适合关注工业AI安全、合规性与模型鲁棒性的研究人员和工程师

将大型语言模型(LLMs)部署为电力系统运行助手虽可提升合规与决策效率,但也引入了基于提示的对抗性攻击风险。本文评估了大模型被越狱的风险,即绕过安全对齐以生成违反监管标准的输出,假设威胁来自授权用户(如操作员)精心设计的恶意提示。测试了三种先进大模型(OpenAI的GPT-4o mini、Google的Gemini 2.0 Flash-Lite、Anthropic的Claude 3.5 Haiku),针对基线、BitBypass和DeepInception三种越狱方法,在源自九项NERC可靠性标准(EOP、TOP、CIP)的场景中进行评估。初步实验中总体攻击成功率(ASR)为33.1%,其中DeepInception效果最佳,达到63.17%。Claude 3.5 Haiku完全免疫(0% ASR),Gemini 2.0 Flash-Lite最脆弱(55.04% ASR),GPT-4o mini中等易受攻击(44.34% ASR)。后续实验通过优化基线与BitBypass攻击中的恶意措辞,使攻击成功率提升至30.6%,证实细微提示调整可显著增强简单攻击的有效性。

原文摘要 · Abstract (English)

The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk of jailbreaking LLMs, i.e., circumventing safety alignments to produce outputs violating regulatory standards, assuming threats from authorized users, such as operators, who craft malicious prompts to elicit non-compliant guidance. Three state-of-the-art LLMs (OpenAI's GPT-4o mini, Google's Gemini 2.0 Flash-Lite, and Anthropic's Claude 3.5 Haiku) were tested against Baseline, BitBypass, and DeepInception jailbreaking methods across scenarios derived from nine NERC Reliability Standards (EOP, TOP, and CIP). In the initial broad experiment, the overall Attack Success Rate (ASR) was 33.1%, with DeepInception proving most effective at 63.17% ASR. Claude 3.5 Haiku exhibited complete resistance (0% ASR), while Gemini 2.0 Flash-Lite was most vulnerable (55.04% ASR) and GPT-4o mini moderately susceptible (44.34% ASR). A follow-up experiment refining malicious wording in Baseline and BitBypass attacks yielded a 30.6% ASR, confirming that subtle prompt adjustments can enhance simpler methods' efficacy.

大模型安全电网智能越狱攻击合规检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。