用LLM当检索评估员,能评对但未必懂理由。
LLMs as Assessors: Right for the Right Reason?
- 让LLM在维基百科数据集上判断相关性并标出关键段落
- 相比人类评估者,LLM在文档级判断上表现接近,但理由匹配度低
- 适合需要快速生成评估数据的研究者,但不可替代人工
近期研究关注使用大语言模型(LLMs)作为人类评估者的替代者,以评价文本/图像处理系统输出的质量。本文聚焦于信息检索(IR)中的标准即兴检索任务,探讨LLM作为相关性评估者的有效性。我们基于INEX倡议创建的维基百科测试集,要求LLM不仅判断文档是否相关,还需标出其认为有用的段落。该数据集的人类评估者也接受了类似指令,需标注所有回应查询信息需求的段落。这使我们不仅能评估LLM在文档层面的判断质量,还能量化其‘因正确理由而正确’的比例。结果表明:尽管LLM表现出色,可显著减少生成高质量基准数据集所需的人工投入,但仍无法完全替代人类评估者。
原文摘要 · Abstract (English)
A good deal of recent research has focused on how Large Language Models (LLMs) may be used as judges in place of humans to evaluate the quality of the output produced by various text / image processing systems. Within this broader context, a number of studies have investigated the specific question of how effectively LLMs can be used as relevance assessors for the standard ad hoc task in Information Retrieval (IR). We extend these studies by looking at additional questions. Most importantly, we use a Wikipedia based test collection created by the INEX initiative, and prompt LLMs to not only judge whether documents are relevant / non-relevant, but to highlight relevant passages in documents that it regards as useful. The human relevance assessors involved in creating this collection were given analogous instructions, i.e., they were asked to highlight all passages within a document that respond to the information need expressed in a query. This enables us to evaluate the quality of LLMs as judges not only at the document level, but to also quantify how often these judges are right for the right reasons. Our observations lead us to reiterate the cautionary note sounded in some earlier studies when it comes to using LLMs as assessors for creating IR datasets: while LLMs are unquestionably promising, and may be used judiciously to subtantially reduce the amount of human involvement required to generate high-quality benchmark datasets, they cannot replace humans as assessors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。