arXiv:2606.20223cs.CVq-bio.QM2026-06中稿 · ICPR 2026 - Comput…

升级版动物识别系统,精准捕捉非洲热带森林中更多物种与复杂环境。

DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests

论文配图:DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests
图 1 · 摘自论文原文
  • 基于生态梯度扩展分类体系至64类,覆盖垂直分层与人类活动区域
  • 跨国家验证准确率达0.86,河岸区物种识别数从4增至9
  • 保留离线流程,适用于野外多场景部署,尤其适合林缘与河岸监测

非洲热带森林的相机陷阱监测正从密闭林冠内部扩展至河岸、空地及公园边缘。现有开放工具中,DeepForestVision 是唯一提供照片与视频一致离线工作流的系统,且先前表现优于其他基线。但其原设计针对地面林内环境,仅35类分类过于粗略,难以应对树栖灵长类、鸟类、半水生动物及家畜等干扰。本文提出 DeepForestVisionV2,将分类体系扩展至64类(61种动物加人类、车辆、空白),针对垂直分层、场景开敞度和人为界面三类部署梯度进行生态驱动优化。该模型基于153万张照片与24万段视频训练,涵盖多国项目数据。评估结合跨国家裁剪照片验证集与三个乌干达视频基准,覆盖目标梯度。在验证集上,准确率0.86,宏平均F1为0.82,平衡准确率为0.81;在部署基准上,尽管任务更难,仍保持或提升准确率,林内视频识别物种数从22增至29,河岸从4增至9;公园边缘准确率由0.62升至0.86,误报率从11降至0。结果表明,DeepForestVisionV2显著提升实地适用性,同时保持跨站点、栖息地与设备设置的鲁棒性。

原文摘要 · Abstract (English)

Camera-trap monitoring in African tropical forests increasingly extends beyond closed-canopy interiors to riverbanks, clearings, and park edges. Among available open tools for African forest camera-trap classification, DeepForestVision is the only one providing a matched offline workflow for both photographs and videos, and previous work showed that it outperformed other available baselines on a comparable benchmark. However, it was designed for closed-canopy, ground-level forest interiors and uses a 35-class prediction space that becomes too coarse when deployments encounter arboreal primates, birds, semi-aquatic taxa, or human-associated confounders such as livestock. We present DeepForestVisionV2, an ecology-driven expansion from 35 to 64 prediction classes (61 animal classes plus human, vehicle, and blank) designed to address three recurrent deployment gradients: vertical stratification, scene openness, and anthropogenic interfaces. DeepForestVisionV2 retains the same offline workflow and is trained on 1,535,010 photographs and 243,354 videos from multi-country African tropical-forest projects. Evaluation combines a cross-country cropped-photo validation set, used to assess robustness across sites and camera-trap settings, with three held-out Uganda video benchmarks spanning the targeted gradients. On the validation set, DeepForestVisionV2 reaches 0.86 accuracy, 0.82 macro-F1, and 0.81 balanced accuracy. On the deployment benchmarks, it preserves or improves baseline accuracy despite its harder classification task, while increasing the number of identified taxa from 22 to 29 in forest-interior videos and from 4 to 9 at riverbanks. In the park-edge use case, it raises accuracy from 0.62 to 0.86 and reduces false alarms from 11 to 0. These results show that DeepForestVisionV2 materially improves field utility while preserving robustness across sites, habitats, and camera-trap settings.

动物识别相机陷阱生态监测多类别分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。