arXiv:2411.08275cs.IRcs.CL2024-11被引 47

用大模型自动判读相关性,效果接近人工,且无需人机协作。

A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look

  • 用开源工具UMBRELA自动生成相关性判断,替代人工。
  • 自动判断与人工判断在nDCG@20、@100和Recall@100上高度一致。
  • 人机协同未提升准确率,人工反而更严格,适合评估研究者。

本文报告了大规模评估(TREC 2024 RAG Track)结果,比较了四种相关性评估方法:一种长期使用的完全人工标准流程,以及三种利用大语言模型不同程度的替代方案,均通过开源工具UMBRELA实现。该设置使我们能对比不同方法生成的系统排名,分析成本与质量的权衡。结果显示,在来自19个团队的77次运行中,由UMBRELA自动生成的相关性判断与完全人工判断在nDCG@20、nDCG@100和Recall@100上的系统排名高度相关。结果表明,自动生成的判断可准确捕捉运行级有效性,替代全人工评估。令人意外的是,引入大模型辅助并未提高与人工判断的相关性,暗示人机协作带来的成本未带来明显收益。整体来看,人工评估者比UMBRELA更严格。本工作验证了大模型在学术TREC式评估中的可行性,为后续研究奠定基础。

原文摘要 · Abstract (English)

The application of large language models to provide relevance assessments presents exciting opportunities to advance information retrieval, natural language processing, and beyond, but to date many unknowns remain. This paper reports on the results of a large-scale evaluation (the TREC 2024 RAG Track) where four different relevance assessment approaches were deployed in situ: the "standard" fully manual process that NIST has implemented for decades and three different alternatives that take advantage of LLMs to different extents using the open-source UMBRELA tool. This setup allows us to correlate system rankings induced by the different approaches to characterize tradeoffs between cost and quality. We find that in terms of nDCG@20, nDCG@100, and Recall@100, system rankings induced by automatically generated relevance assessments from UMBRELA correlate highly with those induced by fully manual assessments across a diverse set of 77 runs from 19 teams. Our results suggest that automatically generated UMBRELA judgments can replace fully manual judgments to accurately capture run-level effectiveness. Surprisingly, we find that LLM assistance does not appear to increase correlation with fully manual assessments, suggesting that costs associated with human-in-the-loop processes do not bring obvious tangible benefits. Overall, human assessors appear to be stricter than UMBRELA in applying relevance criteria. Our work validates the use of LLMs in academic TREC-style evaluations and provides the foundation for future studies.

信息检索大模型评估自动化判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。