研究人类标注差异对模型训练与评估的影响,提出更优的评价指标和训练方法。
Training and Evaluating with Human Label Variation: An Empirical Study
- 基于模糊集理论设计可微分的新评价指标,支持端到端训练。
- 在6个数据集上测试14种方法,发现使用拆分标注或软标签效果最佳。
- 适合关注标注不确定性、需提升模型鲁棒性的研究人员参考。
人类标注差异(HLV)挑战了标准假设中每个样本有唯一真实标签的观点,转而承认人类标注中的自然变异,并以此为基础进行模型训练与评估。尽管已有多种针对HLV的训练方法和评价指标,但其在不同场景下的表现仍不明确。本文提出基于模糊集理论的新评价指标,因其可微分特性,进一步探索将其作为训练目标的可行性。我们在6个HLV数据集上系统测试了14种训练方法和6种评价指标。结果表明,在各类指标下,基于拆分标注或软标签的训练策略均表现最优,优于使用可微分指标作为训练目标的方法。此外,我们提出的软微F1分数是当前最佳的HLV评价指标之一。
原文摘要 · Abstract (English)
Human label variation (HLV) challenges the standard assumption that a labelled instance has a single ground truth, instead embracing the natural variation in human annotation to train and evaluate models. While various training methods and metrics for HLV have been proposed, it is still unclear which methods and metrics perform best in what settings. We propose new evaluation metrics for HLV leveraging fuzzy set theory. Since these new proposed metrics are differentiable, we then in turn experiment with employing these metrics as training objectives. We conduct an extensive study over 6 HLV datasets testing 14 training methods and 6 evaluation metrics. We find that training on either disaggregated annotations or soft labels performs best across metrics, outperforming training using the proposed training objectives with differentiable metrics. We also show that our proposed soft micro F1 score is one of the best metrics for HLV data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。