arXiv:2510.18196cs.CLcs.AI2025-10ACL被引 2

用对比解码减少大模型评分偏差,提升自动评估可靠性。

Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge

  • 提出对比解码方法缓解评分范围偏差问题。
  • 在不同评分范围内平均提升11.7%与人工评分的相关性。
  • 适合需要可信自动评估的文本摘要场景。

大型语言模型常被用作各类应用中的评估者,但其评估结果的可靠性仍存挑战。本文聚焦摘要任务,发现直接评分(即无参考地给出指定范围分数)存在评分范围偏差——模型输出高度依赖预设评分范围,且同一家族模型间也存在类似偏差。为此,提出对比解码方法进行缓解,在不同评分范围下平均实现11.7%的斯皮尔曼相关性相对提升,显著增强与人工判断的一致性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are commonly used as evaluators in various applications, but the reliability of the outcomes remains a challenge. One such challenge is using LLMs-as-judges for direct assessment, i.e., assigning scores from a specified range without any references. Focusing on summarization, we first show that this challenge stems from LLM judge outputs being associated with score range bias, i.e., LLM judge outputs are highly sensitive to pre-defined score ranges. We also show that similar biases exist among models from the same family. We then mitigate this bias through contrastive decoding, achieving up to 11.7% relative improvement on average in Spearman correlation with human judgments across different score ranges.

大模型评估评分偏差对比解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。