大模型安全检测存在系统性盲区,非能力不足而是机制缺陷。
The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability

- 用可度量的检测边界替代二元不可判定,量化安全监控的遗漏程度。
- 实际检测平均仅60.9%准确率,远低于理论最优100%上限。
- 提升效果主要来自信息与检查流程,而非模型本身能力。
安全监控需验证跨租户隔离、沙袋攻击防御和评估感知等2-安全超性质,这些性质仅由两段执行轨迹决定。标准结论为二元不可判定:单段轨迹无法判断。本文将二元判断转为可测量,提出严格界:任意单轨迹监控的平衡准确率为$\tfrac12+\tfrac12\,TV(P_0,P_1)$,将不可判定转化为可分级的可检测性边界,定义了监督缺口——监控器与其理论上限的差距。在闭式表达的泄漏家族中,九个主流大模型监控器在TV=0时表现最优,但随TV增长信号捕获能力骤降;当TV=1时(20行成员检测达100%),其平均准确率仅为60.9%。该缺口主要非因能力不足:明确检查目标可弥补61%的差距,而原控制组仍处于随机水平。该现象贯穿2×2因子设计:假设第二轮运行未执行时监控器仍为随机(50.4%),同一规则在已执行第二轮时则达90.0%;存储预言机无比较过程仅得68.2%。信息与程序缺一不可,二者均非模型能力。在非确定性下,重放轨迹仅在正确投影下遵循闭式k-重放曲线,投影边界揭示困境必然:窄检测漏掉98.6%的偏渠道泄漏,宽检测误报75.7%正常流量,可达准确率随良性扰动率和通道数按$1/(qm)$衰减。最后,两名前沿大模型裁判认证此前版本基准有效,但符号检验发现方向性偏差(p=2.7×10⁻⁵),使我们三个结论失效。超性质基准的构造有效性应机械证明,而非依赖模型审计。
原文摘要 · Abstract (English)
Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hyperproperties, witnessed only by two executions. The standard consequence is a binary impossibility: one trace cannot decide them. We replace the binary with a measurement. A tight bound puts the balanced accuracy of any single-trace monitor at $\tfrac12+\tfrac12\,TV(P_0,P_1)$, turning undecidability into a graded detectability frontier and defining an oversight gap: a monitor's shortfall below it. On a leak family with closed-form $TV$, nine LLM monitors are optimal at $TV=0$ but capture little signal as $TV$ grows; at $TV=1$, where a 20-line membership check scores $100\%$, they average $60.9\%$. That shortfall is mostly not capability: naming what to check closes $61\%$ of it while leaving the $TV=0$ control at chance. The same split runs through a $2{\times}2$ factorial: an imagined second run leaves monitors at chance ($50.4\%$) while the same rule on an executed second run reaches $90.0\%$, and a stored oracle without a comparison procedure yields only $68.2\%$. Information and procedure are each necessary and neither is capability. Under nondeterminism, replay tracks a closed-form $k$-replay curve only under the right projection, and a projection frontier shows the resulting dilemma is forced: narrow misses $98.6\%$ of off-channel leaks, broad flags $75.7\%$ of clean traffic, and attainable accuracy decays like $1/(qm)$ in the benign-variation rate and the channel count. Finally, two frontier LLM judges certified an earlier version of our own benchmark as sound while a sign test found a directional bias ($p=2.7\times10^{-5}$) that invalidated three of our findings. Construction validity for hyperproperty benchmarks should be proved mechanically, not audited by models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。