发现指令模型存在无声错误承诺,不同架构治理能力差异巨大。
Silent Commitment Failure in Instruction-Tuned Language Models: Evidence of Governability Divergence Across Architectures
- 提出'可治理性'概念,衡量错误在输出前是否可检测修正。
- 三模型中两例出现无预警的自信错误输出,另一模型提前57词发出信号。
- 模型治理能力由预训练决定,微调影响极小,适合安全部署研究者参考。
当大语言模型被赋予自主执行工具的权限时,其安全架构依赖一个关键假设:模型错误可在运行时被检测。我们通过实证发现,三个可评估冲突检测的指令跟随模型中有两个无法满足该假设。本文引入‘可治理性’——即模型错误在输出前是否可检测并修正的程度——并证明其在不同模型间存在显著差异。在十二个推理领域测试的六种模型中,三种指令跟随模型中有两种表现出‘无声承诺失败’:自信且流畅的错误输出,但无任何预警信号。剩余模型在贪婪解码下于输出前57个词产生可检测的冲突信号。我们发现基准准确率无法预测可治理性,检测与纠错能力独立变化,相同治理框架在不同模型上效果相反。2×2实验显示,架构间脉冲比相差52倍,而微调仅带来±0.32倍波动,表明可治理性在预训练阶段已固化。本文提出‘检测与纠正矩阵’,将模型-任务组合划分为四类:可治理、仅监控、盲目引导、不可治理。
原文摘要 · Abstract (English)
As large language models are deployed as autonomous agents with tool execution privileges, a critical assumption underpins their security architecture: that model errors are detectable at runtime. We present empirical evidence that this assumption fails for two of three instruction-following models evaluable for conflict detection. We introduce governability -- the degree to which a model's errors are detectable before output commitment and correctable once detected -- and demonstrate it varies dramatically across models. In six models across twelve reasoning domains, two of three instruction-following models exhibited silent commitment failure: confident, fluent, incorrect output with zero warning signal. The remaining model produced a detectable conflict signal 57 tokens before commitment under greedy decoding. We show benchmark accuracy does not predict governability, correction capacity varies independently of detection, and identical governance scaffolds produce opposite effects across models. A 2x2 experiment shows a 52x difference in spike ratio between architectures but only +/-0.32x variation from fine-tuning, suggesting governability is fixed at pretraining. We propose a Detection and Correction Matrix classifying model-task combinations into four regimes: Governable, Monitor Only, Steer Blind, and Ungovernable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。