LLM判官的可信度评分与真假判断高度相关,不能独立看待。
When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

- 通过控制问题正确性测试,发现模型可信度评分与真假判断关联更强。
- 改变答案来源线索后,模型的可信度和真假判断均发生显著变化。
- 提示模型不可将可信度视为独立证据,适合评估其可靠性时使用。
LLM-as-Judge系统可生成多维度评价,如可信度、可靠性与事实性,这些输出常被视作独立证据。本文检验了常见判断对——可信度评分与二元真假分类——是否可分离。在正确性受控的问答任务中,模型的可信度评分与真假判断的对齐程度高于人类行为参考,表明可信度与真假判断的分离性较弱。随后,通过仅改变相同问答对的答案来源线索进行压力测试,发现来源归属变化不仅影响可信度评分,也改变了真假判断结果及基于对数几率的正确侧概率。结果表明,当前LLM-as-Judge协议不应将可信度评分视为独立于真假判断的证据。
原文摘要 · Abstract (English)
LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidence. We test this assumption for a common pair of judgments: trust scoring and binary truth classification. On correctness-controlled QA, LLM judges align trust scores with truth verdicts more tightly than human behavioral reference, suggesting weaker separations between trust and truth judgment. We then apply stress tests by changing only source cues of identical QA between Human and AI. Source attribution shifts not only trust scores but also truth verdicts and logit-derived correct-side probabilities. Results show that current LLM-as-Judge protocols should not treat trust scores as independent evidence for truth judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。