arXiv:2510.00821cs.AI2025-10

通过逻辑一致性检测专家分歧,提升大模型评估的可靠性。

Logical Consistency Between Disagreeing Experts and Its Role in AI Safety

  • 基于专家决策一致与分歧构建逻辑约束,无需标签进行无监督评估。
  • 利用整数线性规划求解可能正确答案集,识别最低评分阈值违规。
  • 适用于大模型作判官时的可信度验证,尤其适合安全敏感场景。

当两名专家在测试中意见不一,我们可判断二者不可能都完全正确;但若完全一致,则所有可能的评判结果都无法排除。本文探讨了这一不对称性,形式化了一种用于分类器的无监督评估逻辑。核心问题是计算与观察到的专家一致与分歧决策逻辑一致的群体评估集合。将专家对齐决策的统计信息作为输入,构建整数空间中的线性规划问题,求解真实标签下可能正确或错误的响应组合。显式逻辑约束(如正确响应数不超过观测响应数)以不等式形式表达,同时引入适用于所有有限测试的普遍线性等式公理。该方法在实践中已成功实现‘零知识警报’系统,能有效检测一个或多个大模型作为裁判时是否违反用户设定的最低评分阈值。

原文摘要 · Abstract (English)

If two experts disagree on a test, we may conclude both cannot be 100 per cent correct. But if they completely agree, no possible evaluation can be excluded. This asymmetry in the utility of agreements versus disagreements is explored here by formalizing a logic of unsupervised evaluation for classifiers. Its core problem is computing the set of group evaluations that are logically consistent with how we observe them agreeing and disagreeing in their decisions. Statistical summaries of their aligned decisions are inputs into a Linear Programming problem in the integer space of possible correct or incorrect responses given true labels. Obvious logical constraints, such as, the number of correct responses cannot exceed the number of observed responses, are inequalities. But in addition, there are axioms, universally applicable linear equalities that apply to all finite tests. The practical and immediate utility of this approach to unsupervised evaluation using only logical consistency is demonstrated by building no-knowledge alarms that can detect when one or more LLMs-as-Judges are violating a minimum grading threshold specified by the user.

AI安全无监督评估逻辑一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。