通过聚类分析标注者意见,提升主观NLP任务的标签聚合效果。
Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks

- 基于标注者共识度进行聚类,捕捉不同观点群体。
- 在40个数据集上,性能优于多数投票和单个标注者建模。
- 多标签与多任务方法更适合聚类后的标注者建模。
标注不一致是构建NLP数据集时的常见现象,也是宝贵的信息来源。尽管多数投票仍是主流标签聚合策略,近期研究尝试建模个体标注者以保留其观点。然而,建模每个标注者成本高,且在各类NLP任务中仍不充分。本文提出一种基于共识的聚类技术,用于建模标注者间的分歧。我们在18种语言、40个数据集上,涵盖情感分析、情绪分类和仇恨言论检测三类主观任务进行了全面实验。评估了四种聚合方法:多数投票、集成、多标签和多任务。结果表明,基于共识的聚类能充分利用标注者视角差异,在主观任务中显著提升分类性能,优于多数投票和个体标注者建模。多标签与多任务方法在建模聚类后的标注者时表现优于集成与多数投票。
原文摘要 · Abstract (English)
Disagreement in annotation is a common phenomenon in the development of NLP datasets and serves as a valuable source of insight. While majority voting remains the dominant strategy for aggregating labels, recent work has explored modeling individual annotators to preserve their perspectives. However, modeling each annotator is resource-intensive and remains underexplored across various NLP tasks. We propose an agreement-based clustering technique to model the disagreement between the annotators. We conduct comprehensive experiments in 40 datasets in 18 typologically diverse languages, covering three subjective NLP tasks: sentiment analysis, emotion classification, and hate speech detection. We evaluate four aggregation approaches: majority vote, ensemble, multi-label, and multitask. The results demonstrate that agreement-based clustering can leverage the full spectrum of annotator perspectives and significantly enhance classification performance in subjective NLP tasks compared to majority voting and individual annotator modeling. Regarding the aggregation approach, the multi-label and multitask approaches are better for modeling clustered annotators than an ensemble and model majority vote.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。