为大模型生成的电网操作提供实时安全审核,防止物理不可行指令执行。
TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions

- 在电网数字孪生中预演每条指令,验证连通性、潮流、发电与负荷切除约束。
- 500次测试中零误放任不安全操作,但模型误差下误接受率升至30.09%。
- 适合电网自动化系统开发者与安全评估人员,保障大模型控制的安全性。
大型语言模型(LLM)辅助的能源管理工具可将自然语言转化为结构化电网指令,但语法正确不代表物理可行。本文提出TwinGridShield,一种与模型无关的运行时授权层,在释放前通过确定性网络孪生对每项提议动作进行评估。原型检查连通性、支路潮流、发电机和负荷切除不变量,并将每次决策记录于哈希链日志中。在匹配模型的IEEE 14节点实验中,配置为以概率p=0.84选择不安全动作的随机提议源,在500次受攻击状态试验中产生421个不安全提议,实际发生率为84.2%,该值反映设定代理特性而非实测大模型提示注入脆弱性。TwinGridShield在这些试验中实现0例不安全释放。由于动作标记与授权使用相同直流潮流模型、系统状态、支路评级和编码约束,此结果验证了实现与编码授权谓词的一致性,而非模型误差下的安全性。主鲁棒性评估引入模型失配:在±20%每母线负荷测量误差下,不安全接受率达5.63%;当实际支路评级比建模值低20%时,不安全接受率达30.09%。
原文摘要 · Abstract (English)
Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a model-independent runtime authorization layer that evaluates each proposed action in a deterministic network twin before release. The prototype checks connectivity, branch-flow, generator, and load-shedding invariants and records each decision in a hash-chained log. A controlled IEEE 14-bus study evaluates single-step switching, redispatch, and load-shedding actions using DC power flow and experimentally assigned branch ratings. In the matched-model experiment, a stochastic proposal source configured to select an unsafe action with probability p=0.84 produced 421 unsafe proposals in 500 attacked-condition trials, a realized rate of 84.2%. This value characterizes the configured surrogate and is not an empirical measurement of LLM prompt-injection susceptibility. TwinGridShield produced 0 unsafe releases in those 500 trials. Because action labeling and authorization used the same DC model, system state, branch ratings, and encoded constraints, this result verifies conformance of the implementation to its encoded authorization predicate rather than safety under model error. The principal robustness evaluation therefore introduces model mismatch. Unsafe acceptance reached 5.63% under bounded +20% and -20% per-bus load-measurement error and 30.09% when actual branch ratings were 20% below modeled ratings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。