用自动识别系统追踪野生大猩猩,助力濒危物种保护
GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- 构建三类野生灵长类视频数据集,支持大猩猩个体识别
- 多帧自监督预训练提升模型对生物特征的捕捉能力
- 适合野生动物保护与生态监测研究者使用
监控极度濒危的西部低地大猩猩目前受限于从海量相机陷阱视频中手动识别个体所需的巨大人力。自动化的主要障碍在于缺乏可用于训练鲁棒深度学习模型的大规模、真实野外视频数据集。为此,我们引入一个综合性基准,包含三个新数据集:Gorilla-SPAC-Wild(迄今最大的野生灵长类重识别视频数据集)、Gorilla-Berlin-Zoo(用于评估跨域识别泛化性)和Gorilla-SPAC-MoT(用于评估相机陷阱视频中的多目标跟踪)。基于这些数据集,我们提出GorillaWatch,一个集成检测、跟踪与重识别的端到端流程。为利用时间信息,我们设计了一种多帧自监督预训练策略,通过轨迹一致性学习领域特异性特征,无需人工标注。为确保科学有效性,采用可微分的AttnLRP验证模型依赖于判别性生物特征而非背景相关性。大量基准测试表明,融合大规模图像骨干网络特征的表现优于专用视频架构。最后,通过将时空约束融入标准聚类,解决无监督种群计数中的过分割问题。所有代码和数据集均公开发布,以推动濒危物种的可扩展、非侵入式监测。
原文摘要 · Abstract (English)
Monitoring critically endangered western lowland gorillas is currently hampered by the immense manual effort required to re-identify individuals from vast archives of camera trap footage. The primary obstacle to automating this process has been the lack of large-scale, "in-the-wild" video datasets suitable for training robust deep learning models. To address this gap, we introduce a comprehensive benchmark with three novel datasets: Gorilla-SPAC-Wild, the largest video dataset for wild primate re-identification to date; Gorilla-Berlin-Zoo, for assessing cross-domain re-identification generalization; and Gorilla-SPAC-MoT, for evaluating multi-object tracking in camera trap footage. Building on these datasets, we present GorillaWatch, an end-to-end pipeline integrating detection, tracking, and re-identification. To exploit temporal information, we introduce a multi-frame self-supervised pretraining strategy that leverages consistency in tracklets to learn domain-specific features without manual labels. To ensure scientific validity, a differentiable adaptation of AttnLRP verifies that our model relies on discriminative biometric traits rather than background correlations. Extensive benchmarking subsequently demonstrates that aggregating features from large-scale image backbones outperforms specialized video architectures. Finally, we address unsupervised population counting by integrating spatiotemporal constraints into standard clustering to mitigate over-segmentation. We publicly release all code and datasets to facilitate scalable, non-invasive monitoring of endangered species
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。