用大模型评估播客相关性,发现人类专家更信大模型而非原有标注。
Revisiting Human-vs-LLM judgments using the TREC Podcast Track
- 在播客片段上测试五种大模型,重评TREC的标注结果
- 高分歧样本中,人类专家更倾向认同大模型判断
- 提醒单一标注者影响评估质量,适合信息检索研究者
使用大语言模型(LLM)进行相关性标注在信息检索领域日益重要。尽管有研究认为大模型与人工判断高度一致,也有研究持相反观点。现有工作多集中于传统文本检索场景。本文聚焦于由音频转录成两分钟片段的TREC 2020和2021播客赛道数据集,采用五种不同大模型重新评估所有查询-片段对,这些对原由TREC评估员标注。此外,针对大模型与TREC评估员分歧最大的小部分样本再次评估,发现人类专家更倾向于与大模型达成一致。结果支持2002年Sormunen的观点:依赖单一评估者会降低用户一致性。
原文摘要 · Abstract (English)
Using large language models (LLMs) to annotate relevance is an increasingly important technique in the information retrieval community. While some studies demonstrate that LLMs can achieve high user agreement with ground truth (human) judgments, other studies have argued for the opposite conclusion. To the best of our knowledge, these studies have primarily focused on classic ad-hoc text search scenarios. In this paper, we conduct an analysis on user agreement between LLM and human experts, and explore the impact disagreement has on system rankings. In contrast to prior studies, we focus on a collection composed of audio files that are transcribed into two-minute segments -- the TREC 2020 and 2021 podcast track. We employ five different LLM models to re-assess all of the query-segment pairs, which were originally annotated by TREC assessors. Furthermore, we re-assess a small subset of pairs where LLM and TREC assessors have the highest disagreement, and found that the human experts tend to agree with LLMs more than with the TREC assessors. Our results reinforce the previous insights of Sormunen in 2002 -- that relying on a single assessor leads to lower user agreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。