用大模型生成评估标签,大幅降低信息检索人工成本
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
- 8个团队用不同大模型生成TREC 2023的判断标签
- 42组自动生成标签可用于分析模型偏差与效果差异
- 适合研究自动化评估、低资源场景或模型对比的学者
使用大语言模型(LLMs)进行相关性评估为信息检索(IR)、自然语言处理(NLP)等领域带来新机遇。通过自动标注可显著减少构建评估数据集所需的人工工作量,尤其适用于知识匮乏的新主题或人力稀缺的低资源场景。本文报告了在SIGIR 2024举办的LLMJudge挑战赛结果,发布并评测了8个国际团队基于8种不同大模型生成的42组TREC 2023深度学习赛道相关性判断标签。这些多样化自动标签有助于研究大模型带来的系统性偏差,探索集成模型的有效性,分析模型与人类评估者间的权衡,并推动自动化评估方法的发展。相关资源已公开:https://llm4eval.github.io/LLMJudge-benchmark/
原文摘要 · Abstract (English)
Using Large Language Models (LLMs) for relevance assessments offers promising opportunities to improve Information Retrieval (IR), Natural Language Processing (NLP), and related fields. Indeed, LLMs hold the promise of allowing IR experimenters to build evaluation collections with a fraction of the manual human labor currently required. This could help with fresh topics on which there is still limited knowledge and could mitigate the challenges of evaluating ranking systems in low-resource scenarios, where it is challenging to find human annotators. Given the fast-paced recent developments in the domain, many questions concerning LLMs as assessors are yet to be answered. Among the aspects that require further investigation, we can list the impact of various components in a relevance judgment generation pipeline, such as the prompt used or the LLM chosen. This paper benchmarks and reports on the results of a large-scale automatic relevance judgment evaluation, the LLMJudge challenge at SIGIR 2024, where different relevance assessment approaches were proposed. In detail, we release and benchmark 42 LLM-generated labels of the TREC 2023 Deep Learning track relevance judgments produced by eight international teams who participated in the challenge. Given their diverse nature, these automatically generated relevance judgments can help the community not only investigate systematic biases caused by LLMs but also explore the effectiveness of ensemble models, analyze the trade-offs between different models and human assessors, and advance methodologies for improving automated evaluation techniques. The released resource is available at the following link: https://llm4eval.github.io/LLMJudge-benchmark/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。