警惕大模型评估的自我强化陷阱,防止误判误导系统优化
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
- 用大模型生成评估结果,可能因反馈循环导致错误结论
- 实证发现大模型评估存在偏差,长期使用会削弱可靠性
- 提出检测机制与协作框架,推动负责任地使用大模型评估
大型语言模型(LLMs)正被用于评估信息检索(IR)系统,替代传统的人工相关性判断。近期研究表明,基于大模型的评估结果常与人工判断一致,引发部分观点认为无需人工评判;但也有研究担忧其判断的可靠性、有效性及长期影响。当大模型生成的信号被用于系统开发和评估时,评估结果可能形成自我强化循环,导致误导性结论。本文分析了大模型评估在系统迭代中虚假成功的情形,指出偏见强化、可复现性差及评估方法不一致等关键风险。为此,提出量化负面影响的测试方法、约束机制,以及融合大模型判断的可复用测试集构建框架。结合学术界与产业界视角,旨在建立大模型在信息检索评估中的规范使用准则。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empirical studies suggest that LLM-based evaluations often align with human judgments, leading some to suggest that human judges may no longer be necessary, while others highlight concerns about judgment reliability, validity, and long-term impact. As IR systems begin incorporating LLM-generated signals, evaluation outcomes risk becoming self-reinforcing, potentially leading to misleading conclusions. This paper examines scenarios where LLM-evaluators may falsely indicate success, particularly when LLM-based judgments influence both system development and evaluation. We highlight key risks, including bias reinforcement, reproducibility challenges, and inconsistencies in assessment methodologies. To address these concerns, we propose tests to quantify adverse effects, guardrails, and a collaborative framework for constructing reusable test collections that integrate LLM judgments responsibly. By providing perspectives from academia and industry, this work aims to establish best practices for the principled use of LLMs in IR evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。