arXiv:2606.20408cs.CRcs.AI2026-06

测试大模型在核电站控制室中抗持续攻击的能力,发现四款模型均易被攻破但漏洞各不相同。

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

论文配图:NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms
图 1 · 摘自论文原文
  • 用模拟核电站环境,让多角色大模型团队对抗四轮自适应攻击
  • 4种主流模型中8.7%至12.1%的攻击会触发关键安全功能失效
  • 不同模型漏洞几乎不重叠,同一防御策略对不同模型效果相反

大型语言模型(LLM)代理正被提议作为关键系统中的监督组件,但其在持续、自适应对抗压力下的鲁棒性仍缺乏评估。我们提出NRT-Bench,一个针对安全关键系统中LLM操作代理的多轮红队测试基准,模拟核电厂控制室场景。五人操作团队各自由可配置的LLM支持,管理由六个关键安全功能(CSFs)控制的系统,攻击者通过四个通道在有限轮次内注入消息,并在每轮后获得反馈。伤害以客观信号衡量:任一CSF丢失即终止运行,归因于引发该事件的消息。在固定攻击配对重放协议下评估四种前沿操作模型,发现自适应多轮攻击可稳定突破安全边界:四模型中,8.7%至12.1%的攻击会致关键安全功能失效。尽管总体失败率相近,但失败案例几乎无重叠——149次会话中,无一次攻击同时攻破所有模型,三分之一至少攻破一个,表明模型漏洞近乎互斥而非嵌套。防御措施的效果高度依赖模型:相同的护栏堆栈或安全顾问代理可能降低某一模型的攻击成功率,却提高另一模型的攻击成功率。我们开源仿真环境、攻击数据集及重放工具,支持可复现的LLM代理安全评估。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as operators of a safety-critical system, instantiated in a simulated nuclear power plant control room. A five-role operator team, each backed by a configurable LLM, runs a plant governed by six critical safety functions (CSFs), while adversaries inject messages over four channels in bounded multi-turn sessions with per-turn feedback. Harm is an objective signal rather than LLM-judged text: a run terminates the moment any CSF is lost, attributed to the causing message. Evaluating four frontier operator models under a fixed-attack paired-replay protocol, we find that adaptive multi-turn attacks reliably push the operator team past a safety limit: across the four models, between 8.7% and 12.1% of attack sessions end with the plant losing a critical safety function. Although the four models look almost equally robust by this aggregate rate, their failures barely overlap: of $149$ sessions, none defeat all four models while a third defeat at least one, so vulnerabilities are nearly disjoint across models rather than nested. The effect of added defences is strongly model-dependent: the same guardrail stack or safety-advisor agent that lowers attack success for one model can raise it for another. We release the simulation venue, attack dataset, and replay tooling for reproducible safety evaluation of LLM agents.

安全评估红队测试大模型代理核能控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。