arXiv:2410.17099cs.CLcs.HC2024-10EMNLP被引 8

让人类和大模型协作,更好合并众包文本答案。

Human-LLM Hybrid Text Answer Aggregation for Crowd Annotations

  • 设计人机协同的多阶段框架,让大模型负责整合
  • 在公开数据集上验证,融合效果优于纯人工或纯模型
  • 适合需要高质量文本标注的众包任务

众包标注的质量是关键问题。答案聚合是重要解决方案,最终的标注结果来自对同一任务多个众包回答的整合,而非原始个体回答。近年来,大语言模型(LLM)在数据标注任务中的能力受到关注。现有研究多聚焦于个体众包工作者的平均表现;少数工作研究了类别标签的聚合以及将LLM用作标签生成者的情况。然而,针对文本答案的聚合场景,以及将LLM作为聚合者角色的研究仍不充分。本文探究了在封闭式众包文本回答聚合场景中,使用大模型作为聚合者的可行性。提出一种人类-大模型混合文本答案聚合方法,基于创作者-聚合者多阶段(CAMS)众包框架。实验基于公开众包数据集进行,结果表明,人类与大模型协同的方法具有显著有效性。

原文摘要 · Abstract (English)

The quality is a crucial issue for crowd annotations. Answer aggregation is an important type of solution. The aggregated answers estimated from multiple crowd answers to the same instance are the eventually collected annotations, rather than the individual crowd answers themselves. Recently, the capability of Large Language Models (LLMs) on data annotation tasks has attracted interest from researchers. Most of the existing studies mainly focus on the average performance of individual crowd workers; several recent works studied the scenarios of aggregation on categorical labels and LLMs used as label creators. However, the scenario of aggregation on text answers and the role of LLMs as aggregators are not yet well-studied. In this paper, we investigate the capability of LLMs as aggregators in the scenario of close-ended crowd text answer aggregation. We propose a human-LLM hybrid text answer aggregation method with a Creator-Aggregator Multi-Stage (CAMS) crowdsourcing framework. We make the experiments based on public crowdsourcing datasets. The results show the effectiveness of our approach based on the collaboration of crowd workers and LLMs.

众包标注大模型聚合人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。