用GPT-4复现27个真实研究任务,发现大模型标注表现差异大。
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
- 用GPT-4复现11个真实研究数据集的27项标注任务
- 大模型在不同任务间表现差异显著,部分偏离人工判断
- 强调必须用人标注作为基准,才能可靠评估自动化工具
自动化文本标注是生成式大语言模型在社交媒体研究中的重要应用。尽管近期研究显示大模型在标注任务上表现优异,但这些研究多基于少量任务且依赖公开基准数据集,存在数据污染风险。本文采用以人为中心的评估框架,使用GPT-4复现11个来自高影响力期刊的计算社会科学论文中密码保护数据集上的27项标注任务。每项任务均将GPT-4标注结果与人工标注真值及基于人工标签微调的监督分类模型进行对比。尽管大模型整体标注质量较高,但其表现存在显著任务间差异,甚至在同一数据集中也表现不一。结果表明,必须建立以人类为中心的工作流和严格评估标准:即使经过提示调优等优化策略,自动化标注仍可能显著偏离人类判断。因此,以人类生成的验证标签为基础,是实现负责任评估的关键。
原文摘要 · Abstract (English)
Automated text annotation is a compelling use case for generative large language models (LLMs) in social media research. Recent work suggests that LLMs can achieve strong performance on annotation tasks; however, these studies evaluate LLMs on a small number of tasks and likely suffer from contamination due to a reliance on public benchmark datasets. Here, we test a human-centered framework for responsibly evaluating artificial intelligence tools used in automated annotation. We use GPT-4 to replicate 27 annotation tasks across 11 password-protected datasets from recently published computational social science articles in high-impact journals. For each task, we compare GPT-4 annotations against human-annotated ground-truth labels and against annotations from separate supervised classification models fine-tuned on human-generated labels. Although the quality of LLM labels is generally high, we find significant variation in LLM performance across tasks, even within datasets. Our findings underscore the importance of a human-centered workflow and careful evaluation standards: Automated annotations significantly diverge from human judgment in numerous scenarios, despite various optimization strategies such as prompt tuning. Grounding automated annotation in validation labels generated by humans is essential for responsible evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。