对比人类与AI生成提示对大模型相关性判断的影响,发现提示差异显著影响结果。
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
- 收集30份提示(15人+15模型),测试其在三种任务中的敏感性
- 不同提示导致大模型与人工标注的吻合度差异达0.26(κ值)
- 揭示提示设计对评估可靠性至关重要,适合评估研究者参考
大型语言模型(LLMs)越来越多地用于信息检索(IR)任务中的相关性判断,其与人工标签的一致性接近人与人之间的共识水平。为评估基于大模型的相关性判断的鲁棒性与可靠性,我们系统研究了提示敏感性的影响。从15位人类专家和15个大模型中收集了三种任务(二元、分级、成对)的相关性评估提示,共获得90个提示。剔除3名人类和3个模型的无效提示后,使用剩余72个提示,由三个不同大模型对TREC深度学习数据集(2020年和2021年)中的文档/查询对进行标注。通过Cohen's κ和成对一致性度量,将大模型生成的标签与TREC官方人工标签进行比较。除了分析提示变化对与人工标签一致性的影响外,还对比了人类与大模型生成的提示,并分析了不同大模型作为裁判的差异。此外,还将人类与大模型生成的提示与Bing及TREC 2024 RAG赛道使用的标准UMBRELA提示进行比较。为支持未来基于大模型的评估研究,所有数据与提示已发布于https://github.com/Narabzad/prompt-sensitivity-relevance-judgements/。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to automate relevance judgments for information retrieval (IR) tasks, often demonstrating agreement with human labels that approaches inter-human agreement. To assess the robustness and reliability of LLM-based relevance judgments, we systematically investigate impact of prompt sensitivity on the task. We collected prompts for relevance assessment from 15 human experts and 15 LLMs across three tasks~ -- ~binary, graded, and pairwise~ -- ~yielding 90 prompts in total. After filtering out unusable prompts from three humans and three LLMs, we employed the remaining 72 prompts with three different LLMs as judges to label document/query pairs from two TREC Deep Learning Datasets (2020 and 2021). We compare LLM-generated labels with TREC official human labels using Cohen's $κ$ and pairwise agreement measures. In addition to investigating the impact of prompt variations on agreement with human labels, we compare human- and LLM-generated prompts and analyze differences among different LLMs as judges. We also compare human- and LLM-generated prompts with the standard UMBRELA prompt used for relevance assessment by Bing and TREC 2024 Retrieval Augmented Generation (RAG) Track. To support future research in LLM-based evaluation, we release all data and prompts at https://github.com/Narabzad/prompt-sensitivity-relevance-judgements/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。