arXiv:2608.23941cs.AI2026-08

研究发现:长审查窗口让模型更排斥而非更精准,短单位更优。

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

论文配图:More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight
图 1 · 摘自论文原文
  • 提出双前缀框架,隔离审查单位影响
  • 审查越长,误拒率同步上升,精准度峰值在1-2步
  • 适合关注安全控制与模型评估的研究者

预执行监督是可信AI管控的核心:由不可靠的大模型监控器在动作执行前进行审查。过度拦截会牺牲可用性,迫使部署方禁用监督机制。每个协议必须明确审查单位——一次调用审查多少动作。现有设计将单位视为既定,其对错误监测器的影响未被测量。自然数据无法分离该因素:审查长度与错误类型、位置共变。仅看召回率会误导:全拒可捕获所有错误。测量需独立变动边界并设置清洁对照。本文提出双前缀框架,满足两者要求。每个真实计划生成一个含一个环境认可错误的前缀,及一个仅在一次写入上不同的清洁副本。在五个嵌套长度下分别判断成对样本,将判断变化归因于审查单位本身。通过预注册的知情度(召回减去误拒)评分。结果显示,审查越长,召回率提升,但误拒率同步上升。六位评委在两个领域中,知情度峰值均出现在1或2个动作;更长窗口使零样本监控器更排斥,而非更判别。重放被隐藏的观测表明失败主要源于观察缺失。安全论证应明示审查单位并附清洁序列数据。本框架是首个受控、预注册的测量工具,从不单独依赖召回率。校准后的短单位可恢复高达0.95的知情度,优于八步审查,且无测试过的标签无关策略能持续超越它。

原文摘要 · Abstract (English)

Pre-execution oversight is core to trusted monitoring in AI control: a fallible LLM monitor vets planned actions before irreversible execution. Over-blocking forfeits usefulness and pressures deployers to disable it. Every protocol must fix a unit of verification: how many actions one call reviews. Existing designs take the unit as given; its effect on fallible monitors is unmeasured. Natural traces cannot isolate it: review length co-varies with error type and position. Catch alone misleads: rejecting everything catches everything. Measuring this needs boundary variation alone and a matched clean control. We introduce the twin-prefix framework, which supplies both. Each gold plan yields a prefix with one injected, environment-accepted error and a clean twin differing in one write. Judging each pair at five nested lengths ties verdict changes to the unit alone. Discrimination is scored by pre-registered informedness, catch minus false rejection. Longer review raises catch; false rejection climbs in lockstep. Informedness peaks at one or two actions for all six judges in both domains: longer windows make zero-shot monitors more rejective, not more discriminative. Replaying withheld observations traces the failure largely to observation deprivation. Safety cases should state the unit and co-report the clean series. Our framework is the first controlled, pre-registered instrument for this choice and never reads catch alone. Our calibrated short unit recovers up to 0.95 informedness over eight-action review, and no tested label-blind policy consistently beats it.

大模型安全审查单位监控评估零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。