arXiv:2510.27106cs.CL2025-10EMNLP被引 49

大模型当评分员,打分结果反复无常。

Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks

  • 用大模型自评生成文本,但多次评分差异大。
  • 不同运行下评分波动显著,可靠性低。
  • 适合关注评估可信度的研究者。

随着自然语言生成(NLG)的广泛应用,其评估变得愈发困难。近期,使用大语言模型(LLM)作为评价工具逐渐流行,因其评分更贴近人类偏好,优于传统的n-gram或嵌入式指标。然而,我们在实验中发现,LLM作为评分员在不同运行间表现出低内部一致性,评分波动大,极端情况下近乎随机,难以真实衡量其判断质量。我们量化了该不一致性在不同NLG任务和基准上的表现,并探讨在遵循严格规范的前提下,是否仍可有效使用LLM评分。

原文摘要 · Abstract (English)

As Natural Language Generation (NLG) continues to be widely adopted, properly assessing it has become quite difficult. Lately, using large language models (LLMs) for evaluating these generations has gained traction, as they tend to align more closely with human preferences than conventional n-gram or embedding-based metrics. In our experiments, we show that LLM judges have low intra-rater reliability in their assigned scores across different runs. This variance makes their ratings inconsistent, almost arbitrary in the worst case, making it difficult to measure how good their judgments actually are. We quantify this inconsistency across different NLG tasks and benchmarks and see if judicious use of LLM judges can still be useful following proper guidelines.

大模型评估评分一致性LLM打分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。