提出统计检验方法,用小样本数据判断LLM能否替代人工标注。
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
- 设计替代标注者检验(alt-test),仅需少量样本即可验证LLM可行性
- 实验显示部分闭源LLM(如GPT-4o)已可超越开源模型表现
- 提供可解释的评估指标,帮助研究者选择最优提示策略
大型语言模型(LLMs)作为标注者、裁判和评估者,在传统由人类完成的任务中被广泛应用。尽管其对研究结果有重要影响,但尚无标准程序来判断是否可用LLM替代人类标注。本文提出一种新的统计方法——替代标注者检验(alt-test),仅需少量标注样本即可合理判定使用LLM的可行性。同时引入一个通用且可解释的度量方式,用于比较不同LLM标注者与裁判的表现。我们构建了涵盖语言与视觉-语言任务的十套数据集,对六种LLM及四种提示技术进行了实验。结果显示,部分闭源模型(如GPT-4o)在某些任务上可超越开源模型,且提示策略显著影响裁判质量。本研究旨在推动更严谨可靠的评估实践。
原文摘要 · Abstract (English)
The "LLM-as-an-annotator" and "LLM-as-a-judge" paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans. LLM annotations are widely used, not only in NLP research but also in fields like medicine, psychology, and social science. Despite their role in shaping study results and insights, there is no standard or rigorous procedure to determine whether LLMs can replace human annotators. In this paper, we propose a novel statistical procedure, the Alternative Annotator Test (alt-test), that requires only a modest subset of annotated examples to justify using LLM annotations. Additionally, we introduce a versatile and interpretable measure for comparing LLM annotators and judges. To demonstrate our procedure, we curated a diverse collection of ten datasets, consisting of language and vision-language tasks, and conducted experiments with six LLMs and four prompting techniques. Our results show that LLMs can sometimes replace humans with closed-source LLMs (such as GPT-4o), outperforming the open-source LLMs we examine, and that prompting techniques yield judges of varying quality. We hope this study encourages more rigorous and reliable practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。