arXiv:2607.27849eess.SYcs.LG2026-07

用规则门控确保大模型控制安全,防止越界操作。

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

论文配图:Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings
图 1 · 摘自论文原文
  • 设计九约束的反事实双分支规则门控,实时拦截违规指令
  • 门控使干扰抑制性能提升16倍,误动作减少至0/10
  • 适合工业控制、安全关键系统中大模型应用的可信部署

一个开放权重的大语言模型每五分钟可生成控制设定点。但工厂仍需要硬性审查:明确的约束条件、日志化的安全裕度,以及在监管层执行前的允许/拒绝决策。本文在不改变监管层的前提下,引入基于规则的分叉双分支反事实门控(包含九个固定约束),实现这一硬检查。在Skogestad列A上,对比了仅PID(C0)、线性MPC(C1)、无门控代理(C2)和有门控代理(C3)四种模式,在相同工况、种子与水平闭合目标(M_D, M_B)下进行测试。结果显著:非额定工况下的目标追踪性能,代理优于帕累托调优的线性MPC(C2/C1 IAE比为0.361,上置信区间);扰动抑制能力在16点网格上反转达16.03倍(点估计值10.18),无门控代理已不适用。该门控将规格偏离吸引子压缩为有限偏移(d≈-1.4;P95单元的IAE从11.5降至0.77)。一句提示词修复即可从源头消除吸引子(6/10降至0/10;仅敏感性改善,非新标题级成果)。在250单元统计测试中,534/590次干预属于规范边界几何:当操作规范恰好位于安全极限时,正常操作者被禁用,而异常行为者仅被限制;其中318次阻断了实际有害提案。所有结论均针对DeepSeek-V4-Flash单列模型得出。第二组测试(NVIDIA Nemotron-3-Super)保持扰动抑制失败区域与工厂侧故障地理特征不变,性能幅度与协议可用性仍受模型影响,且强性能细胞仅为幸存者(非确认)。迁移意义在于双分支结构、约束包络与设定点接口,而非测量另一类工厂。

原文摘要 · Abstract (English)

An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, and an admit/block decision before the regulatory layer moves. This paper puts that check in a rule-based forked-twin counterfactual gate (nine pinned constraints) and leaves the regulatory layer unchanged. On Skogestad's Column A the ladder is PID-only (C0), linear MPC (C1), ungated agent (C2), and gated agent (C3) under one contract: identical level closure (M_D, M_B), scenarios, and seeds; C2/C3 share the linear-MPC backend. The split is not subtle. Off-nominal target acquisition: the agent beats Pareto-tuned linear MPC in the strong band (C2/C1 IAE ratio 0.361 at the upper CI). Disturbance rejection on the same 16-point grid inverts by 16.03 at the upper CI (10.18 at the point estimate), where an ungated LLM supervisor does not belong. The gate compresses a specification-abandonment attractor into a bounded offset (d approx. -1.4; P95 cell IAE 11.5 to 0.77). A one-line prompt fix removes the attractor at source (6/10 to 0/10; sensitivity only, not a new headline). In a 250-cell statistical pass, 534 of 590 gate interventions are spec-on-bound geometry: the operating specification sits on a safety limit, so a well-behaved OP becomes inoperable while misbehaving ones are only contained; 318 blocks still correct actively harmful proposals. Headlines are single-column and model-conditional on DeepSeek-V4-Flash. A second-family sweep (NVIDIA Nemotron-3-Super) keeps the disturbance-rejection fails band and plant-side failure geography; magnitudes and protocol operability stay model-conditional, and Super target-acquisition strong cells are survivors only (not confirmation). Transfer means twin, constraint envelope, and setpoint interface, not a second plant class measured here.

大模型控制安全门控工业自动化反事实机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。