别用大模型生成搜索评估的相关性判断,效果不可靠。
Don't Use LLMs to Make Relevance Judgments
- 用大模型自动标注相关性判断,结果与人工标注差异大
- 在TREC评测中,大模型判断准确率远低于专业人员
- 适合需要高精度评估的科研人员参考,避免误用
TREC风格测试集的相关性判断工作复杂且成本高昂,通常需六名承包商投入2至4周时间,并经过培训与监控。为确保判断记录正确高效,还需开发专用软件。近年来,大语言模型(LLMs)能以自然语言提示生成类人流畅文本,激发了信息检索(IR)研究者对其用于相关性判断过程的设想。在ACM SIGIR 2024会议上,'LLM4Eval'研讨会为此提供了平台,并组织数据挑战活动,要求参赛者复现Thomas等人(arXiv:2408.08896, arXiv:2309.10621)在深度学习赛道上的相关性判断。本文基于该研讨会主旨演讲撰写而成,核心结论明确:不建议使用大语言模型生成TREC式评估的相关性判断。
原文摘要 · Abstract (English)
Making the relevance judgments for a TREC-style test collection can be complex and expensive. A typical TREC track usually involves a team of six contractors working for 2-4 weeks. Those contractors need to be trained and monitored. Software has to be written to support recording relevance judgments correctly and efficiently. The recent advent of large language models that produce astoundingly human-like flowing text output in response to a natural language prompt has inspired IR researchers to wonder how those models might be used in the relevance judgment collection process. At the ACM SIGIR 2024 conference, a workshop ``LLM4Eval'' provided a venue for this work, and featured a data challenge activity where participants reproduced TREC deep learning track judgments, as was done by Thomas et al (arXiv:2408.08896, arXiv:2309.10621). I was asked to give a keynote at the workshop, and this paper presents that keynote in article form. The bottom-line-up-front message is, don't use LLMs to create relevance judgments for TREC-style evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。