arXiv:2608.30789cs.CV2026-08

用志愿者标注的不确定标签训练模型,反而提升难样本识别准确率。

Camera trap classification with deep learning under ground truth uncertainty

论文配图:Camera trap classification with deep learning under ground truth uncertainty
图 1 · 摘自论文原文
  • 在标注有分歧的数据上训练,模型对难图识别更准
  • 预训练ImageNet可减少训练轮次,降低计算成本
  • 适合生态图像分类中利用志愿者数据的实践者

监督式深度学习可快速处理生态图像数据,但依赖昂贵的人工标注。因此,训练标签常来自志愿公民科学项目,但志愿者间意见不一致导致‘真实标签’存在不确定性。我们使用两个包含相机陷阱图像及志愿者与专家分类结果的数据集,研究了在更高标签不确定性下训练的影响。结果显示,整体测试准确率提高,尤其在志愿者难以判断的图像上表现更优;物种级准确率普遍提升,但在跨数据集泛化上未见改善。使用ImageNet预训练增强了不确定性带来的收益,减少了训练轮次;进一步在其他相机陷阱图像上预训练可降低计算成本,但未提升准确率。在类别不平衡数据下,仍观察到增加标签不确定性对整体准确率的益处,尤其对难样本。类别不平衡提升了常见物种准确率,降低了稀有物种准确率,且误判模式更接近志愿者错误。研究结果对多标签生态图像分类具有启示,建议在训练中引入适度标签分歧,并使用通用图像预训练模型,以更好整合人类与深度学习分类结果。

原文摘要 · Abstract (English)

Supervised deep learning methods enable the rapid processing of ecological image data, but depend on a costly annotation process. Consequently, training labels are commonly derived from volunteer citizen science projects. However, disagreement among volunteers introduces uncertainty in the "ground truth" data that are assumed to be correct for model training and validation. Using two datasets containing camera trap images with associated volunteer and expert classifications, we investigated the effects of training under higher ground truth uncertainty. We observed improved overall test accuracy, particularly for images that were more difficult for volunteers. Species-level accuracy also generally improved, but generalisation to a different dataset did not. The benefits of ground truth uncertainty were enhanced by pre-training on ImageNet. Pre-training also reduced the number of training epochs required; further reductions in computational cost, but not gains in accuracy, resulted from additional pre-training on other camera trap images. With unbalanced training data, we still observed a clear benefit of increased ground truth uncertainty for overall accuracy, especially on difficult images. Class imbalance improved accuracy for common species, reduced rare species accuracy, and changed patterns of misclassification to more closely resemble mistakes made by volunteers. Our findings have implications for applying deep learning across ecological image types with multiple labels. Practitioners can improve accuracy, especially on difficult examples, by including moderate levels of label disagreement during training and using models pre-trained on general image data. In addition to improving the use of citizen science-derived labels in model training, our study suggests avenues for more effectively integrating human and deep learning classifications in combined workflows. (abridged)

深度学习生态监测标注不确定性志愿者数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。