用大模型辅助情感标注,提升效率与质量
Rethinking Emotion Annotations in the Era of Large Language Models
- 用GPT-4辅助标注,降低人工成本
- 实验显示其标注质量接近人类水平
- 适合需要大规模标注的AI情感研究者
当前情感计算系统严重依赖人工标注的情感数据集,但人工标注成本高、易受设计影响且难以质检,因情绪具有主观性。与此同时,大语言模型(LLMs)在自然语言理解任务中表现卓越,成为文本标注的潜在工具。本文以GPT-4为例,分析其在情感标注中的应用复杂性。实验表明,GPT-4在人类评估中获得高分,表现优于以往仅以人工标注为标准的研究。然而,仍观察到人类与GPT-4在情感感知上的差异,凸显人类输入的重要性。为此,我们探索了两种将GPT-4融入情感标注流程的方法,证明其可有效识别低质量标签、减轻人工负担,并提升下游模型的学习性能与效率。结果表明,大模型有望成为辅助人工标注的有力工具,推动新型情感标注实践的发展。
原文摘要 · Abstract (English)
Modern affective computing systems rely heavily on datasets with human-annotated emotion labels, for training and evaluation. However, human annotations are expensive to obtain, sensitive to study design, and difficult to quality control, because of the subjective nature of emotions. Meanwhile, Large Language Models (LLMs) have shown remarkable performance on many Natural Language Understanding tasks, emerging as a promising tool for text annotation. In this work, we analyze the complexities of emotion annotation in the context of LLMs, focusing on GPT-4 as a leading model. In our experiments, GPT-4 achieves high ratings in a human evaluation study, painting a more positive picture than previous work, in which human labels served as the only ground truth. On the other hand, we observe differences between human and GPT-4 emotion perception, underscoring the importance of human input in annotation studies. To harness GPT-4's strength while preserving human perspective, we explore two ways of integrating GPT-4 into emotion annotation pipelines, showing its potential to flag low-quality labels, reduce the workload of human annotators, and improve downstream model learning performance and efficiency. Together, our findings highlight opportunities for new emotion labeling practices and suggest the use of LLMs as a promising tool to aid human annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。