arXiv:2503.19092cs.IRcs.AI2025-03被引 36

揭示大模型在信息检索评估中的相互偏倚,警示评估可靠性

Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation

  • 用实验检验大模型排名器、助手与评判器的互动影响
  • 发现评判模型严重偏向排名模型,难分辨细微性能差异
  • 适合关注大模型评估可信度的研究者和系统设计者

大语言模型(LLMs)正广泛用于信息检索(IR),驱动排序、评估与内容生成。然而其组件间的交互可能引入偏差。本文整合现有研究并提出新实验设计,探究基于大模型的排序器与助手对大模型评判器的影响。首次提供实证证据:大模型评判器显著偏向大模型排序器。同时发现,大模型评判器难以识别系统间细微性能差异。与部分先前研究相反,初步结果未显示对人工智能生成内容的偏见。这些发现凸显需以整体视角审视大模型驱动的信息生态。为此,本文提出初步指导原则与研究议程,保障大模型在信息检索评估中的可靠使用。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly integral to information retrieval (IR), powering ranking, evaluation, and AI-assisted content creation. This widespread adoption necessitates a critical examination of potential biases arising from the interplay between these LLM-based components. This paper synthesizes existing research and presents novel experiment designs that explore how LLM-based rankers and assistants influence LLM-based judges. We provide the first empirical evidence of LLM judges exhibiting significant bias towards LLM-based rankers. Furthermore, we observe limitations in LLM judges' ability to discern subtle system performance differences. Contrary to some previous findings, our preliminary study does not find evidence of bias against AI-generated content. These results highlight the need for a more holistic view of the LLM-driven information ecosystem. To this end, we offer initial guidelines and a research agenda to ensure the reliable use of LLMs in IR evaluation.

大模型评估信息检索偏倚分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。