对比作者自述与第三方标注,发现后者难以准确捕捉真实情绪。
Can Third-parties Read Our Emotions?
- 用真人实验对比作者自评与第三方标注的情绪
- 第三方标注普遍存在偏差,大模型表现优于人工
- 相似背景的标注者或加入作者信息能提升标注质量
旨在从文本中推断作者内在状态(如情绪、观点)的自然语言处理任务,通常依赖第三方标注的数据集。然而,第三方能否准确捕捉作者私有状态这一假设尚未被充分检验。本研究通过人类受试者实验,直接比较了第三方标注与作者自述情绪标签在情绪识别任务中的表现。结果表明,无论是人工标注还是大语言模型(LLMs)生成的标注,均存在显著局限性,难以忠实反映作者的真实状态。但总体上,大语言模型的表现优于人类标注者。我们进一步探索提升标注质量的方法:发现作者与标注者之间的背景相似性可提升人工标注效果;在提示中加入作者人口统计信息,能带来微弱但统计显著的性能提升。本文提出一个评估第三方标注局限性的框架,并呼吁改进标注实践,以更准确地建模作者的内在状态。
原文摘要 · Abstract (English)
Natural Language Processing tasks that aim to infer an author's private states, e.g., emotions and opinions, from their written text, typically rely on datasets annotated by third-party annotators. However, the assumption that third-party annotators can accurately capture authors' private states remains largely unexamined. In this study, we present human subjects experiments on emotion recognition tasks that directly compare third-party annotations with first-party (author-provided) emotion labels. Our findings reveal significant limitations in third-party annotations-whether provided by human annotators or large language models (LLMs)-in faithfully representing authors' private states. However, LLMs outperform human annotators nearly across the board. We further explore methods to improve third-party annotation quality. We find that demographic similarity between first-party authors and third-party human annotators enhances annotation performance. While incorporating first-party demographic information into prompts leads to a marginal but statistically significant improvement in LLMs' performance. We introduce a framework for evaluating the limitations of third-party annotations and call for refined annotation practices to accurately represent and model authors' private states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。