arXiv:2409.12218cs.CLcs.LG2024-09AAAI被引 10

用自一致性检测标注者可靠性,提升数据质量

ARTICLE: Annotator Reliability Through In-Context Learning

  • 通过上下文学习实现标注一致性自检
  • 在两个仇恨言论数据集上验证效果优于传统方法
  • 适合需要高质量人工标注的NLP任务

确保NLP训练与评估数据中标注者的质量是关键问题。情感分析和仇恨言论检测等任务具有内在主观性,传统质量评估方法难以区分因工作不认真导致的分歧与真实观点差异。为在保持多样视角的同时保证一致性,我们提出 exttt{ARTICLE}——一种基于上下文学习(ICL)的标注质量估计框架,通过自一致性判断标注可靠性。我们在两个仇恨言论数据集上使用多种大语言模型进行了评估,并与传统方法对比。结果表明, exttt{ARTICLE}可作为识别可靠标注者的一种稳健方法,从而提升数据质量。

原文摘要 · Abstract (English)

Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to distinguish disagreement due to poor work from that due to differences of opinions between sincere annotators. With the goal of increasing diverse perspectives in annotation while ensuring consistency, we propose \texttt{ARTICLE}, an in-context learning (ICL) framework to estimate annotation quality through self-consistency. We evaluate this framework on two offensive speech datasets using multiple LLMs and compare its performance with traditional methods. Our findings indicate that \texttt{ARTICLE} can be used as a robust method for identifying reliable annotators, hence improving data quality.

标注质量大模型自一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。