测试大模型能否像人类一样隐藏在标注者群体中
Can We Hide Machines in the Crowd? Quantifying Equivalence in LLM-in-the-loop Annotation Tasks
- 用统计方法检验大模型与人类标注者的决策是否等价
- 在MovieLens数据集上大模型与人类无法区分(p=0.004)
- 可提前用少量数据判断大模型是否适合大规模标注
当前大语言模型在文本标注中的评估多关注输出正确性,通常通过标准指标将模型生成标签与人工标注的“真实标签”对比。本文超越单一有效性评估,探索人类与大模型标注决策在个体间的统计可比性。不将大模型视为单纯标注系统,而是将其视为可模仿人类主观判断的替代标注机制。为此,提出基于Krippendorff's α、配对自助法和两样本等效性t检验(TOST)的统计评估方法,用于检测大模型是否能在群体中不可分辨。该方法应用于MovieLens 100K和PolitiFact两个数据集,结果显示在MovieLens 100K上大模型与人类无统计差异(p = 0.004),但在PolitiFact上则存在差异(p = 0.155),表明任务特性影响等效性。该方法还支持基于小样本人类数据进行早期评估,以决定大模型是否适用于特定场景的大规模标注。
原文摘要 · Abstract (English)
Many evaluations of large language models (LLMs) in text annotation focus primarily on the correctness of the output, typically comparing model-generated labels to human-annotated ``ground truth'' using standard performance metrics. In contrast, our study moves beyond effectiveness alone. We aim to explore how labeling decisions -- by both humans and LLMs -- can be statistically evaluated across individuals. Rather than treating LLMs purely as annotation systems, we approach LLMs as an alternative annotation mechanism that may be capable of mimicking the subjective judgments made by humans. To assess this, we develop a statistical evaluation method based on Krippendorff's $α$, paired bootstrapping, and the Two One-Sided t-Tests (TOST) equivalence test procedure. This evaluation method tests whether an LLM can blend into a group of human annotators without being distinguishable. We apply this approach to two datasets -- MovieLens 100K and PolitiFact -- and find that the LLM is statistically indistinguishable from a human annotator in the former ($p = 0.004$), but not in the latter ($p = 0.155$), highlighting task-dependent differences. It also enables early evaluation on a small sample of human data to inform whether LLMs are suitable for large-scale annotation in a given application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。