发现Tobacco3482数据集11.7%标签错误,影响模型评估可靠性
Label Errors in the Tobacco3482 Dataset
- 人工检查发现11.7%样本标签不当,16.7%样本应有多个标签
- 顶级模型35%错误直接源于数据标签噪声
- 提醒研究者慎用有标注问题的基准数据集
Tobacco3482是广泛使用的文档分类基准数据集。然而,我们对整个数据集进行人工检查后发现存在广泛的本体论问题,尤其是大量标注标签错误。我们制定了数据标签规范,发现11.7%的数据集样本标注不当,应为未知标签或修正标签,且16.7%的样本具有多个有效标签。随后分析了表现最好的模型,发现其35%的错误可直接归因于这些标签问题,凸显使用噪声标签数据集作为基准的固有问题。补充材料(包括标注和代码)可在https://github.com/gordon-lim/tobacco3482-mistakes/ 获取。
原文摘要 · Abstract (English)
Tobacco3482 is a widely used document classification benchmark dataset. However, our manual inspection of the entire dataset uncovers widespread ontological issues, especially large amounts of annotation label problems in the dataset. We establish data label guidelines and find that 11.7% of the dataset is improperly annotated and should either have an unknown label or a corrected label, and 16.7% of samples in the dataset have multiple valid labels. We then analyze the mistakes of a top-performing model and find that 35% of the model's mistakes can be directly attributed to these label issues, highlighting the inherent problems with using a noisily labeled dataset as a benchmark. Supplementary material, including dataset annotations and code, is available at https://github.com/gordon-lim/tobacco3482-mistakes/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。