arXiv:2411.13212cs.IR2024-11被引 6

LLM生成的检索评估不靠谱,会歪曲顶尖系统的真实排名和统计差异。

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation

  • 用大模型自动打分,但会扭曲顶尖检索系统的实际排序
  • 相比人工评判,大模型导致虚假显著性差异比例极高
  • 提醒研究者谨慎使用,尤其关注系统间微小差距的验证

离线检索系统评估依赖于测试集,包含文档、查询和相关性判断。尽管测试集是信息检索研究的基础,其构建需大量人工标注。大语言模型(LLMs)被广泛视为自动相关性评估的工具,近期研究表明其评分与人工判断具有较高排名相关性。然而,这种相关性在大规模实验中虽有帮助,对顶尖系统间的比较却信息有限。更重要的是,它忽略了大模型评分是否能准确反映系统间相对于人工判断的统计显著差异。本文研究大模型评分在保留顶尖系统排序差异及配对显著性检验方面的能力。结果表明,大模型评分在排名顶尖系统时存在不公平性,且统计差异的误报率极高。本工作推进了对基于大模型评分的可靠性评估研究,期望为后续开发更可靠的自动相关性评估方法提供基础。

原文摘要 · Abstract (English)

Offline evaluation of search systems depends on test collections. These benchmarks provide the researchers with a corpus of documents, topics and relevance judgements indicating which documents are relevant for each topic. While test collections are an integral part of Information Retrieval (IR) research, their creation involves significant efforts in manual annotation. Large language models (LLMs) are gaining much attention as tools for automatic relevance assessment. Recent research has shown that LLM-based assessments yield high systems ranking correlation with human-made judgements. These correlations are helpful in large-scale experiments but less informative if we want to focus on top-performing systems. Moreover, these correlations ignore whether and how LLM-based judgements impact the statistically significant differences among systems with respect to human assessments. In this work, we look at how LLM-generated judgements preserve ranking differences among top-performing systems and also how they preserve pairwise significance evaluation as human judgements. Our results show that LLM-based judgements are unfair at ranking top-performing systems. Moreover, we observe an exceedingly high rate of false positives regarding statistical differences. Our work represents a step forward in the evaluation of the reliability of using LLMs-based judgements for IR evaluation. We hope this will serve as a basis for other researchers to develop more reliable models for automatic relevance assessment.

信息检索大模型评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。