arXiv:2602.17170cs.IR2026-02中稿 · SIGIR 2026被引 4

LLM评估相关性时普遍高估,可能误导信息检索评价。

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment

  • 对比不同模型和评估方式,发现LLM普遍存在高分倾向
  • 长文本和关键词表面相似性会显著提升评分,即使内容无关
  • 适合关注LLM评估可靠性的研究者与评测框架设计者

人工相关性评估耗时且认知负荷高,限制了信息检索评估的可扩展性。为此,越来越多研究尝试使用大语言模型(LLMs)作为人类评判者的替代。然而,LLM生成的相关性判断是否可靠、稳定且严谨,仍是一个开放问题。本文系统研究了不同模型架构、评估范式(点对点与成对)及段落修改策略下的LLM过评行为。结果表明,模型对不符合信息需求的段落普遍给出过高评分,且信心度很高,显示出系统性偏差而非随机波动。受控实验进一步显示,LLM评分对段落长度和表面词汇线索极为敏感。这些发现警示:直接用LLM替代人类评估存在风险,亟需建立诊断性评估框架。代码与数据已公开。

原文摘要 · Abstract (English)

Human relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using large language models (LLMs) as proxies for human judges. However, it remains an open question whether LLM-based relevance judgments are reliable, stable, and rigorous enough to match humans for relevance assessment. In this work, we conduct a study of \textit{overrating behavior} in LLM-based relevance judgments across model backbones, evaluation paradigms (pointwise and pairwise), and passage modification strategies. We show that models consistently assign inflated relevance scores -- often with high confidence -- to passages that do not genuinely satisfy the underlying information need, revealing a system-wide bias rather than random fluctuations in judgment. Furthermore, controlled experiments show that LLM-based relevance judgments can be highly sensitive to passage length and surface-level lexical cues. These results raise concerns about the usage of LLMs as drop-in replacements for human relevance assessors, and highlight the urgent need for careful diagnostic evaluation frameworks when applying LLMs for relevance assessments. Our code and results are publicly available.

LLM评估相关性判断过评问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。