面向细粒度分类的噪声数据集,用于测试鲁棒学习与标签修正方法。
Noisy Ostracods: A Fine-Grained, Imbalanced Real-World Dataset for Benchmarking Robust Machine Learning and Label Correction Methods
- 构建了包含多种真实噪声的海星类微化石分类数据集
- 数据集存在高度不均衡(不平衡因子22429)和5.58%的标签噪声
- 适合研究抗噪算法、生物分类与弱监督学习的研究者
我们提出Noisy Ostracods数据集,用于甲壳类软体动物的属与种级分类任务,包含71,466个样本,其中5.58%在属级别被估计为存在潜在问题。该数据集涵盖开放集噪声及伪类别问题,即标注者将应属于已有类别的样本误分至新伪类。数据集具有极高的不均衡性,不平衡因子ρ=22429。现有鲁棒学习方法在该数据集上表现不佳,相较于交叉熵训练未见显著提升;噪声检测方法在错误标签识别率上也逊于朴素交叉验证集成。这些结果表明,细粒度、高度不均衡与复杂噪声特性对当前抗噪算法构成严峻挑战。通过公开发布该数据集及其评估协议,旨在推动针对真实世界多样噪声的鲁棒机器学习方法研究。数据集与评测流程可从https://github.com/H-Jamieu/Noisy_ostracods获取。
原文摘要 · Abstract (English)
We present the Noisy Ostracods, a noisy dataset for genus and species classification of crustacean ostracods with specialists' annotations. Over the 71466 specimens collected, 5.58% of them are estimated to be noisy (possibly problematic) at genus level. The dataset is created to addressing a real-world challenge: creating a clean fine-grained taxonomy dataset. The Noisy Ostracods dataset has diverse noises from multiple sources. Firstly, the noise is open-set, including new classes discovered during curation that were not part of the original annotation. The dataset has pseudo-classes, where annotators misclassified samples that should belong to an existing class into a new pseudo-class. The Noisy Ostracods dataset is highly imbalanced with a imbalance factor $ρ$ = 22429. This presents a unique challenge for robust machine learning methods, as existing approaches have not been extensively evaluated on fine-grained classification tasks with such diverse real-world noise. Initial experiments using current robust learning techniques have not yielded significant performance improvements on the Noisy Ostracods dataset compared to cross-entropy training on the raw, noisy data. On the other hand, noise detection methods have underperformed in error hit rate compared to naive cross-validation ensembling for identifying problematic labels. These findings suggest that the fine-grained, imbalanced nature, and complex noise characteristics of the dataset present considerable challenges for existing noise-robust algorithms. By openly releasing the Noisy Ostracods dataset, our goal is to encourage further research into the development of noise-resilient machine learning methods capable of effectively handling diverse, real-world noise in fine-grained classification tasks. The dataset, along with its evaluation protocols, can be accessed at https://github.com/H-Jamieu/Noisy_ostracods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。