高效评估标注员可靠性,提升数据质量与模型性能。
Efficient Annotator Reliability Assessment with EffiARA
- 基于可靠性的软标签聚合与样本加权,优化标注数据
- 通过替换不可靠标注员,提升整体标注一致性
- 提供易用的开源工具包与网页界面,支持全流程管理
数据标注是机器学习流程中的关键环节,但成本高、耗时长。随着基于Transformer的文档级标注兴起,尚缺乏标准化框架。EffiARA是首个覆盖标注全流程的系统,从任务资源评估、数据集构建到个体标注员与数据集整体可靠性分析。已有两项研究验证其有效性:一是通过标注员可靠性进行软标签聚合与样本加权,提升分类性能;二是通过移除并替换不可靠标注员,提高整体标注一致性。本文发布EffiARA Python包及配套网页工具,提供可视化操作界面。代码与工具已开源,项目地址为https://github.com/MiniEggz/EffiARA,网页工具访问地址为https://effiara.gate.ac.uk。
原文摘要 · Abstract (English)
Data annotation is an essential component of the machine learning pipeline; it is also a costly and time-consuming process. With the introduction of transformer-based models, annotation at the document level is increasingly popular; however, there is no standard framework for structuring such tasks. The EffiARA annotation framework is, to our knowledge, the first project to support the whole annotation pipeline, from understanding the resources required for an annotation task to compiling the annotated dataset and gaining insights into the reliability of individual annotators as well as the dataset as a whole. The framework's efficacy is supported by two previous studies: one improving classification performance through annotator-reliability-based soft-label aggregation and sample weighting, and the other increasing the overall agreement among annotators through removing identifying and replacing an unreliable annotator. This work introduces the EffiARA Python package and its accompanying webtool, which provides an accessible graphical user interface for the system. We open-source the EffiARA Python package at https://github.com/MiniEggz/EffiARA and the webtool is publicly accessible at https://effiara.gate.ac.uk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。