arXiv:2602.07673cs.CL2026-02

LLM评估摘要时更偏爱机器生成内容,尤其当与人工摘要相似度低时。

Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation

  • 以ROUGE/BLEU重叠度为指标,分析LLM对摘要的评判偏差
  • 随着与人工摘要相似度下降,LLM更倾向给机器生成摘要高分
  • 多数模型存在此偏差,即使自身也生成过类似文本仍无法避免

大型语言模型(LLM)常被用作摘要评价工具,因其能更好捕捉语义、推理能力强且抗改写。然而,现有研究发现其存在长度、顺序等偏差,且易受对抗性提示影响。本文针对摘要任务中与人工摘要的重叠程度,系统分析了9个参数量在10亿至120亿之间的LLM(含Gemma 3和LLaMA 3变体)的评判偏差。结果表明:当待评摘要与人工参考摘要的重叠度(以ROUGE和BLEU衡量)降低时,多数模型反而更偏好机器生成摘要,该现象在除一个模型外的所有测试模型中均成立,且不受模型自身生成倾向的影响。此外,即使重叠度很低,模型仍难以准确判断摘要质量,说明单纯依赖重叠度比较不足以支撑可靠的LLM评判。

原文摘要 · Abstract (English)

Large language model (LLM) judges have often been used alongside traditional, algorithm-based metrics for tasks like summarization because they better capture semantic information, are better at reasoning, and are more robust to paraphrasing. However, LLM judges show biases for length and order among others, and are vulnerable to various adversarial input prompts. While recent studies have looked into these biases, few have analyzed them at a more granular level in relation to a well-defined overlap metric. In this work we provide an LLM judge bias analysis as a function of overlap with human-written responses in the domain of summarization. We test 9 recent LLMs with parameter counts ranging from 1 billion to 12 billion, including variants of Gemma 3 and LLaMA 3. We find that LLM judges increasingly prefer summaries generated by other LLMs over those written by humans as the similarities (as measured by ROUGE and BLEU) between the judged summaries decrease, and this pattern extends to all but one model tested, and exists regardless of the models' own position biases. Additionally, we find that models struggle to judge even summaries with limited overlaps, suggesting that LLM-as-a-judge in the summary domain should rely on techniques beyond a simple comparison.

LLM评估摘要评价重叠偏差模型偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。