对比四类标注者在德语情感分析中的质量,指导低资源语言数据构建。
Annotation Quality in Aspect-Based Sentiment Analysis: A Case Study Comparing Experts, Students, Crowdworkers, and Large Language Model

- 用专家重标注建立金标准,评估学生、众包、LLM的标注质量。
- 专家标注一致性和下游模型性能最优,众包和LLM次之。
- 为低资源NLP场景提供标注效率与可靠性权衡参考。
方面级情感分析(ABSA)通过识别文本中特定方面的情感,实现细粒度意见分析。尽管英文领域的研究已很成熟,但德语等其他语言的研究仍受限于高质量标注数据的缺乏。本文研究不同标注来源对德语ABSA发展的影响。通过专家重标注现有数据集建立金标准,用于评估学生、众包工作者、大语言模型(LLMs)及专家的标注质量。采用标注一致性(IAA)衡量标注可靠性,并分析其对下游子任务(如方面类别情感分析ACSA和目标方面情感检测TASD)模型性能的影响。实验涵盖BERT、T5和LLaMA-based等前沿方法,包括微调和指令提示下的上下文学习。结果揭示了标注可靠性与效率之间的权衡,为低资源自然语言处理场景的数据构建提供了实用指导。
原文摘要 · Abstract (English)
Aspect-Based Sentiment Analysis (ABSA) enables fine-grained opinion analysis by identifying sentiments toward specific aspects or targets within a text. While ABSA has been widely studied for English, research on other languages such as German remains limited, largely due to the lack of high-quality annotated datasets. This paper examines how different annotation sources influence the development of German ABSA. To this end, an existing dataset is re-annotated by experts to establish a ground truth, which serves as a reference for evaluating annotations produced by students, crowdworkers, Large Language Models (LLMs), and experts. Annotation quality is compared using Inter-Annotator Agreement (IAA) and its impact on downstream model performance for different ABSA subtasks. The evaluation focuses on Aspect Category Sentiment Analysis (ACSA) and Target Aspect Sentiment Detection (TASD). We apply State-of-the-Art (SOTA) methods for ABSA, including BERT-, T5-, and LLaMA-based approaches to assess performance differences, spanning fine-tuning and in-context learning with instruction prompts. The findings provide practical insights into trade-offs between annotation reliability and efficiency, offering guidance for dataset construction in under-resourced Natural Language Processing (NLP) scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。