arXiv:2601.18891cs.CV2026-01被引 1

用弱监督预训练提升北极驯鹿检测精度,解决背景复杂、目标小等难题。

Weakly supervised framework for wildlife detection and counting in challenging Arctic environments: a case study on caribou (Rangifer tarandus)

  • 基于空/非空标签的弱监督图像块预训练,增强模型鲁棒性
  • 多群落数据集上2017年与2019年测试集F1达93.7%和92.6%
  • 适合在标注少的野生动物监测中使用,尤其适用于大范围航拍图

近年来北极地区驯鹿数量持续下降,亟需可扩展且准确的监测方法以支持科学保护决策。人工解读航拍影像耗时且易出错,亟需自动可靠的检测技术。然而,由于背景高度异质、大片空旷区域(类别不平衡)、目标小或被遮挡、密度与尺度变化大,自动检测极具挑战。为此,提出一种基于检测网络架构的弱监督图像块级预训练框架,以提升检测模型(HerdNet)的鲁棒性。该检测数据集涵盖阿拉斯加五个驯鹿群。通过学习空/非空标签,该方法在多群落影像(2017年)及独立年份(2019年)测试集上分别取得F1值93.7%和92.6%,显著优于从通用权重初始化的HerdNet。迁移至检测任务后,弱监督预训练在正样本块(F1: 92.6%/93.5% vs. 89.3%/88.6%)与整图计数(F1: 95.5%/93.3% vs. 91.5%/90.4%)上均表现更优。主要局限在于动物状背景杂波导致的误报,以及低密度遮挡引发的漏检。总体表明,在标注数据有限时,利用粗标签进行弱监督预训练,可获得媲美通用权重初始化的性能。

原文摘要 · Abstract (English)

Caribou across the Arctic has declined in recent decades, motivating scalable and accurate monitoring approaches to guide evidence-based conservation actions and policy decisions. Manual interpretation from this imagery is labor-intensive and error-prone, underscoring the need for automatic and reliable detection across varying scenes. Yet, such automatic detection is challenging due to severe background heterogeneity, dominant empty terrain (class imbalance), small or occluded targets, and wide variation in density and scale. To make the detection model (HerdNet) more robust to these challenges, a weakly supervised patch-level pretraining based on a detection network's architecture is proposed. The detection dataset includes five caribou herds distributed across Alaska. By learning from empty vs. non-empty labels in this dataset, the approach produces early weakly supervised knowledge for enhanced detection compared to HerdNet, which is initialized from generic weights. Accordingly, the patch-based pretrain network attained high accuracy on multi-herd imagery (2017) and on an independent year's (2019) test sets (F1: 93.7%/92.6%, respectively), enabling reliable mapping of regions containing animals to facilitate manual counting on large aerial imagery. Transferred to detection, initialization from weakly supervised pretraining yielded consistent gains over ImageNet weights on both positive patches (F1: 92.6%/93.5% vs. 89.3%/88.6%), and full-image counting (F1: 95.5%/93.3% vs. 91.5%/90.4%). Remaining limitations are false positives from animal-like background clutter and false negatives related to low animal density occlusions. Overall, pretraining on coarse labels prior to detection makes it possible to rely on weakly-supervised pretrained weights even when labeled data are limited, achieving results comparable to generic-weight initialization.

野生动物检测弱监督学习北极生态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。