提出检测大模型判官组共识风险的方法,避免错误一致导致误判。
JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

- 通过校准探针分析判官间假阴性相关性和共识提升,评估共识风险
- 在虚假事实数据上,高风险场景下97%的误接受被成功拦截
- 适合需要低成本判官但又怕集体盲点的自动验证系统使用
越来越多低成本的大语言模型判官用于接受或升级判断。在事实性场景中,因多个无参考判官达成一致而接受某说法,可能隐藏风险:一致可能是共享的假阴性盲区,而非独立证据。我们提出JuryProbe,一种针对无参考事实性判官组的实证共识风险诊断工具,配合基于校准的路由策略。JuryProbe利用标注校准探针,通过仅假阴性(FN-only)判官相关性和假共识提升来估计共识风险;当风险高时,将原判官组的多数接受决策转至使用可信参考的同一组判官进行验证。在经过审计的FEVER数据污染测试中,无参考判官组表现出显著假阴性相关性(FN-only相关系数0.402和0.368;共识提升3.13倍和18.13倍),而在可信参考下的理想诊断中,统一错误共识降至零。在高风险标记场景中,路由策略本质上等同于对所有无参考多数接受进行接地验证(在34/34分割中验证有效):改进来自接受条件下的接地,而诊断决定是否激活。固定预设规则在合成、基准作者和科学类数据中成功标记8-10/10的高风险分割,负控组为0/10,停用可减少28%的参考获取,同时仅增加0.004的误接受率。即使在弱BM25检索条件下,误接受减少仍持续存在,但覆盖度下降;过期停用标签需定期重校准。JuryProbe不提供正式风险保证,也无法在自然判官组上建立可靠停用机制;其核心贡献是提供一种实证性的高风险面板错误依赖诊断。
原文摘要 · Abstract (English)
Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independent evidence. We introduce JuryProbe, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy. JuryProbe estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift; when flagged high-risk, reference-free majority accepts are routed to the same judges with trusted references. On audited FEVER corruptions, reference-free panels show correlated false negatives (FN-only correlations 0.402 and 0.368; lifts 3.13x and 18.13x), while unanimous false consensus drops to zero under a trusted-reference best-case diagnostic on both minimal-pair and non-minimal-pair evidence. In flagged settings, the routed policy is by construction equivalent to grounding every reference-free majority accept (verified in 34/34 splits): improvement comes from accept-conditioned grounding, while the diagnostic determines whether to activate it. A fixed, pre-specified rule flags 8-10 of 10 splits across synthetic, benchmark-authored, and scientific families and 0 of 10 on a negative control, where standing down avoids 28% of reference acquisitions at a 0.004 increase in false accepts. False-accept reduction persists under weak BM25 retrieval at substantial coverage cost, while stale stand-down labels require periodic recalibration. JuryProbe provides no formal risk guarantee and does not establish reliable stand-down on natural panels; its supported contribution is an empirical diagnostic of high-risk panel error dependence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。