arXiv:2503.01691cs.CVcs.LG2025-03NeurIPS被引 6

构建昆虫新物种识别基准数据集,推动生态监测中的开放集识别研究

Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring

  • 构建大规模细粒度昆虫数据集,支持跨区域未知物种检测评估
  • 38种开放集识别算法对比显示后处理方法仍是强基线
  • 利用辅助数据可提升数据稀缺区域的物种发现能力

全球生物多样性正以空前速度衰退,但绝大多数物种及其种群变化仍不为人知,估计地球90%的物种尚未被记录。机器学习为长期、大规模生物多样性监测提供了新可能,尤其在图像驱动的细粒度物种分类方面。然而,现有算法通常无法识别训练中未见的类别,即开放集识别(OSR)问题,限制了其在昆虫等高度多样且研究不足类群中的应用。为此,我们提出Open-Insect——一个大规模、细粒度的数据集,用于评估不同地理区域下未知物种检测的性能,涵盖不同难度。我们在三类方法上评测了38种OSR算法:后处理型、训练时正则化型和使用辅助数据训练型,发现简单后处理方法仍是有效基线。同时证明,通过引入辅助数据可显著提升数据稀缺区域的物种发现能力。研究成果为计算机视觉在生物多样性监测与物种发现中的发展提供重要指导。

原文摘要 · Abstract (English)

Global biodiversity is declining at an unprecedented rate, yet little information is known about most species and how their populations are changing. Indeed, some 90% of Earth's species are estimated to be completely unknown. Machine learning has recently emerged as a promising tool to facilitate long-term, large-scale biodiversity monitoring, including algorithms for fine-grained classification of species from images. However, such algorithms typically are not designed to detect examples from categories unseen during training -- the problem of open-set recognition (OSR) -- limiting their applicability for highly diverse, poorly studied taxa such as insects. To address this gap, we introduce Open-Insect, a large-scale, fine-grained dataset to evaluate unknown species detection across different geographic regions with varying difficulty. We benchmark 38 OSR algorithms across three categories: post-hoc, training-time regularization, and training with auxiliary data, finding that simple post-hoc approaches remain a strong baseline. We also demonstrate how to leverage auxiliary data to improve species discovery in regions with limited data. Our results provide insights to guide the development of computer vision methods for biodiversity monitoring and species discovery.

开放集识别生物多样性昆虫识别细粒度分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。