arXiv:2412.17156cs.IR2024-12被引 61

LLM评估不能替代人工评估,存在严重局限

LLM-based relevance assessment still can't replace human relevance assessment

  • 用伪造系统证明评估指标可被人为抬高
  • 模拟发现依赖LLM会扭曲系统排名顺序
  • 适合关注评估可信度的研究者阅读

大语言模型(LLM)在信息检索相关性评估中的应用备受关注,近期研究声称其判断可媲美人类。尤其基于TREC 2024数据,Upadhyay等人宣称如Umbrela系统等LLM评估可完全替代传统人工评估。本文对此提出质疑:首先,挑战该结论的数据是否真实支持其主张,尤其当测试集本应作为未来创新基准;其次,构造一刻意利用自动评估指标的系统,使其获得虚高分数但未提升检索质量;第三,通过模拟假设所有系统均采用Umbrela作为最终重排序器的情形,分析肯德尔tau相关性,揭示依赖LLM评估将导致系统排名失真。理论层面还存在诸多挑战,包括LLM固有的自我中心倾向、对LLM评估指标的过拟合风险,以及未来LLM性能可能因此退化的隐患。这些都需解决才能使LLM评估成为人类评估的可行替代。

原文摘要 · Abstract (English)

The use of large language models (LLMs) for relevance assessment in information retrieval has gained significant attention, with recent studies suggesting that LLM-based judgments provide comparable evaluations to human judgments. Notably, based on TREC 2024 data, Upadhyay et al make a bold claim that LLM-based relevance assessments, such as those generated by the Umbrela system, can fully replace traditional human relevance assessments in TREC-style evaluations. This paper critically examines this claim, highlighting practical and theoretical limitations that undermine the validity of this conclusion. First, we question whether the evidence provided by Upadhyay et al. genuinely supports their claim, particularly when the test collection is intended to serve as a benchmark for future research innovations.Second, we submit a system deliberately crafted to exploit automatic evaluation metrics, demonstrating that it can achieve artificially inflated scores without truly improving retrieval quality. Third, we simulate the consequences of circularity by analyzing Kendall's tau correlations under the hypothetical scenario in which all systems adopt Umbrela as a final-stage re-ranker, illustrating how reliance on LLM-based assessments can distort system rankings. Theoretical challenges - including the inherent narcissism of LLMs, the risk of overfitting to LLM-based metrics, and the potential degradation of future LLM performance - that must be addressed before LLM-based relevance assessments can be considered a viable replacement for human judgments.

LLM评估信息检索可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。