arXiv:2509.08593cs.AIstat.ML2025-09被引 1

用逻辑一致性检测大模型评委是否失准,无需真实答案。

No-Knowledge Alarms for Misaligned LLMs-as-Judges

  • 通过分析多个模型评委的分歧,构建可验证的评价约束
  • 在无真实答案情况下,能100%准确发现评委能力不达标
  • 适合用于评估模型自评系统可靠性,防止虚假可信

当使用大语言模型作为评委来评估其他模型的复杂决策时,谁来监督这些评委?在缺乏专家真实判断且无法完全信任的情况下,监控链条会无限延伸。缓解评估不确定性的一种方法是利用不同评委间判断的逻辑一致性。通过观察多个大模型评委在评分其他模型时的同意与分歧情况,可以推导出其评分能力的唯一可能评估结果。例如,若两个评委对第三个模型的完成任务判断不一致,则二者不可能都100%正确。这一逻辑可形式化为一个整数响应计数空间中的线性规划问题。本文据此提出无知识警报机制,可在不产生误报的前提下,检测出至少一个评委未满足用户指定的评分能力要求。

原文摘要 · Abstract (English)

If we use LLMs as judges to evaluate the complex decisions of other LLMs, who or what monitors the judges? Infinite monitoring chains are inevitable whenever we do not know the ground truth of the decisions by experts and we do not want to trust them. One way to ameliorate our evaluation uncertainty is to exploit the use of logical consistency between disagreeing experts. By observing how LLM judges agree and disagree while grading other LLMs, we can compute the only possible evaluations of their grading ability. For example, if two LLM judges disagree on which tasks a third one completed correctly, they cannot both be 100\% correct in their judgments. This logic can be formalized as a Linear Programming problem in the space of integer response counts for any finite test. We use it here to develop no-knowledge alarms for misaligned LLM judges. The alarms can detect, with no false positives, that at least one member or more of an ensemble of judges are violating a user specified grading ability requirement.

模型评估逻辑一致性无真值检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。