arXiv:2507.06537cs.CV2025-07

用主动学习选关键照片,30%数据达到全量训练效果

A model-agnostic active learning approach for animal detection from camera traps

  • 不依赖具体模型,从图像和目标两个层面选最值的样本
  • 仅用30%标注数据,就达到甚至超过全量数据性能
  • 适合野外监测中数据量大、标注成本高的场景

智能数据选择在数据驱动的机器学习中日益重要。主动学习通过筛选最具信息量的样本,实现高效模型训练。野生动物相机陷阱采集的数据量巨大,标注与模型训练耗费大量人力。将主动学习应用于优化标注数据量,可显著提升自动化野生动物监测与保护效率。然而,现有方法需访问完整模型,限制了适用性。本文提出一种无需依赖模型的主动学习方法,用于相机陷阱动物检测。该方法在目标级和图像级同时融合不确定性与多样性,指导样本选择。我们在基准动物数据集上验证该方法。实验表明,仅使用30%由本方法选出的训练数据,即可使当前最优动物检测器达到或超越使用全部训练数据的性能。

原文摘要 · Abstract (English)

Smart data selection is becoming increasingly important in data-driven machine learning. Active learning offers a promising solution by allowing machine learning models to be effectively trained with optimal data including the most informative samples from large datasets. Wildlife data captured by camera traps are excessive in volume, requiring tremendous effort in data labelling and animal detection models training. Therefore, applying active learning to optimise the amount of labelled data would be a great aid in enabling automated wildlife monitoring and conservation. However, existing active learning techniques require that a machine learning model (i.e., an object detector) be fully accessible, limiting the applicability of the techniques. In this paper, we propose a model-agnostic active learning approach for detection of animals captured by camera traps. Our approach integrates uncertainty and diversity quantities of samples at both the object-based and image-based levels into the active learning sample selection process. We validate our approach in a benchmark animal dataset. Experimental results demonstrate that, using only 30% of the training data selected by our approach, a state-of-the-art animal detector can achieve a performance of equal or greater than that with the use of the complete training dataset.

主动学习动物检测相机陷阱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。