LLM做题评卷时会偏信自己知识,无视给定答案,导致评分失准。
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- 用错误答案替换标准答案,制造模型知识与参考答案冲突
- 多模型测试显示评分可靠性大幅下降,最低降至32%
- 现有提示缓解策略无效,暴露LLM评卷的根本缺陷
尽管大语言模型(LLMs)越来越多地被用作问答(QA)等依赖参考答案的自动评估工具,但其对所提供参考答案的遵循能力仍不清楚。我们发现了一种关键失败模式:当参考答案与模型参数化知识冲突时,评分变得不可靠,严重降低评估准确性。为此,我们引入一个受控的参考答案替换框架,通过将正确答案替换为错误实体,并构建原始与替换参考答案及其对应候选答案的多样化配对。令人惊讶的是,多种判别模型在替换参考答案下评分可靠性急剧下降。实证表明,这种脆弱性源于模型过度依赖自身知识,导致在冲突时忽略给定参考答案。此外,常见基于提示的缓解策略未能有效改善该问题,凸显了以LLM为判官评估方法的根本局限性,亟需强化对参考答案遵循的协议设计。
原文摘要 · Abstract (English)
While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about their ability to adhere to a provided reference. We identify a critical failure mode of such reference-based LLM QA evaluation: when the provided reference conflicts with the judge model's parametric knowledge, the resulting scores become unreliable, substantially degrading evaluation fidelity. To study this phenomenon systematically, we introduce a controlled swapped-reference QA framework that induces reference-belief conflicts. Specifically, we replace the reference answer with an incorrect entity and construct diverse pairings of original and swapped references with correspondingly aligned candidate answers. Surprisingly, grading reliability drops sharply under swapped references across a broad set of judge models. We empirically show that this vulnerability is driven by judges' over-reliance on parametric knowledge, leading judges to disregard the given reference under conflict. Finally, we find that this failure persists under common prompt-based mitigation strategies, highlighting a fundamental limitation of LLM-as-a-judge evaluation and motivating reference-based protocols that enforce stronger adherence to the provided reference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。