arXiv:2606.07834cs.SEcs.AI2026-06

大模型判官在混杂证据下易误判,本文揭示其承诺偏差并提出双通道控制方案。

Cherry-pick Override: Unsafe Directional Commitment in LLM Judges under Mixed Evidence

  • 设计双通道验证机制,用结构化证据与置信度分离决策与授权
  • 在AVeriTeC数据集上,84%混杂证据案例出现错误方向承诺
  • 提出外部控制层,可有效防止大模型在矛盾证据中强行站队

大模型判官在处理同时存在支持与反驳证据的命题时,常因系统性承诺机制而产生不安全的方向性判断。当任务规范允许‘冲突’为非方向性结论时,若模型仍输出‘支持’或‘反对’,即构成未经授权的方向承诺,称为樱桃挑选覆盖(Cherry-pick Override, CCO)。本文基于显式任务契约,采用同分母诊断协议、匹配覆盖率自助法及苹果对苹果随机否决零假设,在AVeriTeC的冲突子集(N_C = 150)上发现:三选项判官在超过84%的混杂证据命题中返回方向性结论;三人多数投票在AVeriTeC上将方向性承诺率提升至0.887(95% CI [+0.013, +0.080]),但未在VitaminC-Mixed上复现。常见单通道修复手段(如类型词汇、面板聚合、置信阈值、仅验证者过滤)均留下残余失败:面板聚合在48%的CCO案例中压制了单个判官的‘冲突’异议;模型在纯支持/反驳数据上校准良好(ECE = 0.07),置信度无法区分正确方向承诺与CCO;验证者作为分类器几乎使纯证据准确率减半。最小双通道参考探测器达到任一单通道无法企及的性能点;在随机否决零假设下,其向‘冲突’的提升在AVeriTeC上具有结构性显著性(实证p < 1/2001),在VitaminC-Mixed上趋势相同但较弱,表明为选择性而非幅度效应。因此,我们主张引入外部承诺控制层,将判断生成与授权分离,以结构证据与置信度为正交通道,以‘不承诺’作为路由控制状态。

原文摘要 · Abstract (English)

LLM judges increasingly turn verdicts into system commitments. Under mixed evidence (claims with both supporting and refuting sources) this is unsafe: when the schema exposes CONFLICTING as the authorized non-directional verdict, returning SUPPORTS/REFUTES is an unauthorized directional commitment, a failure we name Cherry-pick Override (CCO). We define CCO under an explicit task contract and report it with a same-denominator diagnostic protocol paired with matched-coverage bootstrap and an apples-to-apples random-veto null. On AVeriTeC's Conflicting subset (N_C = 150), three-option judges return a directional verdict on more than 84% of mixed-evidence claims; under the typed schema, three-judge majority voting amplifies direction-on-conflict on AVeriTeC (0.887 vs. 0.840; 95% CI [+0.013, +0.080]) but does not replicate on VitaminC-Mixed. Walking an intervention ladder of common single-channel fixes (typed vocabulary, panel aggregation, confidence thresholding, validator-only filtering), each leaves a distinct residual failure: panel aggregation suppresses single-judge CONFLICTING dissent in 48% of CCO cases; the panel is well-calibrated for direction (ECE = 0.07 on pure-S/R) so confidence cannot operationally separate CCO from correct directional commits; validator-as-classifier nearly halves pure-evidence accuracy. A minimal two-channel reference probe reaches operating points neither single channel reaches; under the random-veto null its promotion to CONFLICTING is structurally targeted on AVeriTeC (empirical p < 1/2001) and weaker but in the same direction on VitaminC-Mixed, a selectivity result rather than a magnitude one. We argue for an external commitment-control layer that separates verdict generation from commitment authorization, using structural evidence and confidence as orthogonal channels and NO-COMMIT as a routed controller state.

大模型评估判官机制承诺偏差双通道

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。