测试发现生成式模型标注存在系统性偏差,与人工标注差异大。
Are generative AI text annotations systematically biased?
- 用多种大模型和提示词对五个概念进行标注对比
- 模型间相似度高于模型与人工标注的相似度,偏差明显
- 即使F1分数高,仍导致下游结果显著不同
本文通过概念复制方式,评估生成式大语言模型(GLLM)在标注政治内容、互动性、理性、不文明言论和意识形态等五个概念时的偏差。实验使用Llama3.1:8b、Llama3.3:70b、GPT4o、Qwen2.5:72b四种模型,搭配五种提示词。结果显示,尽管各模型在F1分数上表现良好,但其标注结果在出现频率上与人工标注存在显著差异,且模型间重合度高于与人工标注的重合度,表明存在系统性偏差。此外,仅靠F1分数无法反映实际偏差程度,该偏差会引发下游分析结果的实质性改变。
原文摘要 · Abstract (English)
This paper investigates bias in GLLM annotations by conceptually replicating manual annotations of Boukes (2024). Using various GLLMs (Llama3.1:8b, Llama3.3:70b, GPT4o, Qwen2.5:72b) in combination with five different prompts for five concepts (political content, interactivity, rationality, incivility, and ideology). We find GLLMs perform adequate in terms of F1 scores, but differ from manual annotations in terms of prevalence, yield substantively different downstream results, and display systematic bias in that they overlap more with each other than with manual annotations. Differences in F1 scores fail to account for the degree of bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。