提出软标签标注框架,更真实反映情绪标注的主观差异。
Quality and Agreement in Multilabel Emotion Annotation: A Case Study and Evaluation Framework

- 用投票占比构建软标签,保留标注者意见分布
- 软监督使模型预测更贴近真实标注差异
- 适合需要处理主观性强任务的研究者参考
情绪标注具有高度主观性,但多数NLP流程仍假设存在‘黄金标签’,通常通过多数投票生成,并将标注者差异视为噪声。本文通过多标签情绪标注案例研究,探讨标注行为与聚合方式对一致性估计及下游分类器的影响。不将分歧简化为单一标签,而是采用软投票份额标签(含强度加权变体),并使用阈值指标(宏/微-F1)和概率对齐(伯努利交叉熵 SoftBCE)进行评估,辅以数据驱动的分歧诊断。在不同标注模式下,我们发现分歧具有结构性,会在模型行为中留下可测量痕迹:硬标签可能最大化F1,而软监督则使预测更好地反映实际标注者变异与不确定性。结果为多标签情绪数据集的设计、聚合与评估提供实用指导。
原文摘要 · Abstract (English)
Emotion annotation is inherently subjective, yet most NLP pipelines still assume "gold" labels, typically produced by majority voting, and treat annotator variation as noise. In this paper, we present a multilabel emotion annotation case study and use it to examine how annotator behavior and aggregation choices affect both agreement estimates and downstream emotion classifiers. Rather than collapsing disagreement into a single label, we represent targets as soft vote-share labels (including an intensity-weighted variant) and evaluate models using both thresholded metrics (macro-/micro-F1) and probabilistic alignment (Bernoulli cross-entropy SoftBCE), alongside data-derived disagreement diagnostics. Across annotation regimes, we show that disagreement is structured and leaves measurable traces in model behavior: hard labels may maximize F1 metrics, while soft supervision yields predictions that better reflect empirical annotator variance and uncertainty. Our results provide practical guidance for designing, aggregating, and evaluating multilabel emotion datasets when multiple interpretations are plausible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。