arXiv:2411.10074cs.CVq-bio.PE2024-11被引 3

通过置信度筛选提升图像自动标注准确率,实现95%以上高精度标注。

Improving the accuracy of automated labeling of specimen images datasets via a confidence-based process

  • 根据模型置信度设定阈值,拒绝低可信标签以提高整体准确率。
  • 86%初始准确率可提升至95%(舍弃40%标签)或99%(舍弃65%标签)。
  • 适用于生态学研究,支持按需权衡准确率与标注覆盖率。

过去三十年自然历史标本数字化催生了海量标本图像与元数据。为提升数据价值,需进一步标注性状信息,深度学习中的卷积神经网络(CNN)等方法有望大幅减少人工标注需求,但当前自动标注准确率通常仅80%-85%,难以可靠使用。本文提出一种基于置信度的筛选方法:通过评估模型对标签的置信度,并设定用户自定义阈值剔除低置信度结果,显著提升准确率。实验表明,初始准确率为86%的模型,通过选择更高置信度阈值,可实现超过95%准确率(舍弃约40%标签)或超99%准确率(舍弃约65%标签)。该方法灵活适配不同研究对准确率与覆盖范围的需求,使自动标注从不可用变为实用工具。我们利用该方法标注了超过60万份标本的生殖状态,分析揭示了未被充分研究的相关性,同时与已知趋势一致。相关数据集将公开共享,供生态学家按需开展研究。

原文摘要 · Abstract (English)

The digitization of natural history collections over the past three decades has unlocked a treasure trove of specimen imagery and metadata. There is great interest in making this data more useful by further labeling it with additional trait data, and modern deep learning machine learning techniques utilizing convolutional neural nets (CNNs) and similar networks show particular promise to reduce the amount of required manual labeling by human experts, making the process much faster and less expensive. However, in most cases, the accuracy of these approaches is too low for reliable utilization of the automatic labeling, typically in the range of 80-85% accuracy. In this paper, we present and validate an approach that can greatly improve this accuracy, essentially by examining the confidence that the network has in the generated label as well as utilizing a user-defined threshold to reject labels that fall below a chosen level. We demonstrate that a naive model that produced 86% initial accuracy can achieve improved performance - over 95% accuracy (rejecting about 40% of the labels) or over 99% accuracy (rejecting about 65%) by selecting higher confidence thresholds. This gives flexibility to adapt existing models to the statistical requirements of various types of research and has the potential to move these automatic labeling approaches from being unusably inaccurate to being an invaluable new tool. After validating the approach in a number of ways, we annotate the reproductive state of a large dataset of over 600,000 herbarium specimens. The analysis of the results points at under-investigated correlations as well as general alignment with known trends. By sharing this new dataset alongside this work, we want to allow ecologists to gather insights for their own research questions, at their chosen point of accuracy/coverage trade-off.

自动标注置信度筛选生态学图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。