用自然语言识别空拍影像中的多物种,跨环境泛化能力强。
OpenWildlife: Open-Vocabulary Multi-Species Wildlife Detector for Geographically-Diverse Aerial Imagery
- 基于语言感知嵌入和改进的Grounding-DINO框架,支持自然语言指定物种
- 在7个新物种数据集上达0.597 mAP50,微调后最高0.981 mAP50
- 搜索算法仅查33%图像就覆盖95%物种,适合大规模生物多样性监测
我们提出OpenWildlife(OW),一个面向多样化空拍影像的开放词汇野生动物检测器。现有自动化方法在特定场景表现良好,但受限于分类覆盖范围与固定架构,难以跨物种和环境泛化。相比之下,OW利用语言感知嵌入和改进的Grounding-DINO框架,可基于自然语言输入识别陆地与海洋环境中多种物种。在15个数据集上训练,微调后在多数数据集上达到最高0.981 mAP50;在7个包含新物种的数据集上仍保持0.597 mAP50。此外,我们设计了一种结合k近邻与广度优先搜索的高效检索算法,仅分析33%图像即可捕获超过95%物种。为保障可复现性,我们公开源代码与数据集划分,使OW成为全球生物多样性评估中灵活且低成本的解决方案。
原文摘要 · Abstract (English)
We introduce OpenWildlife (OW), an open-vocabulary wildlife detector designed for multi-species identification in diverse aerial imagery. While existing automated methods perform well in specific settings, they often struggle to generalize across different species and environments due to limited taxonomic coverage and rigid model architectures. In contrast, OW leverages language-aware embeddings and a novel adaptation of the Grounding-DINO framework, enabling it to identify species specified through natural language inputs across both terrestrial and marine environments. Trained on 15 datasets, OW outperforms most existing methods, achieving up to \textbf{0.981} mAP50 with fine-tuning and \textbf{0.597} mAP50 on seven datasets featuring novel species. Additionally, we introduce an efficient search algorithm that combines k-nearest neighbors and breadth-first search to prioritize areas where social species are likely to be found. This approach captures over \textbf{95\%} of species while exploring only \textbf{33\%} of the available images. To support reproducibility, we publicly release our source code and dataset splits, establishing OW as a flexible, cost-effective solution for global biodiversity assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。