用自监督学习从相机陷阱视频中训练出强鲁棒的黑猩猩面部嵌入模型
Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
- 基于DINOv2框架,从无标签视频中自动挖掘人脸图像进行训练
- 在Bossou等挑战性数据集上,开放集重识别性能超越有监督基线
- 适用于大规模野生动物种群监测,无需人工标注身份信息
相机陷阱正革新野生动物监测,捕获海量视觉数据;然而,个体动物的手动识别仍是主要瓶颈。本研究提出一种完全自监督的方法,从无标签相机陷阱视频中学习鲁棒的黑猩猩面部嵌入。利用DINOv2框架,在自动挖掘的人脸裁剪图像上训练视觉变换器,无需身份标签。该方法在复杂基准如Bossou上展现出强大的开放集重识别性能,优于有监督基线,且训练全程未使用标注数据。本工作凸显了自监督学习在生物多样性监测中的潜力,为可扩展、非侵入式种群研究开辟新路径。
原文摘要 · Abstract (English)
Camera traps are revolutionising wildlife monitoring by capturing vast amounts of visual data; however, the manual identification of individual animals remains a significant bottleneck. This study introduces a fully self-supervised approach to learning robust chimpanzee face embeddings from unlabeled camera-trap footage. Leveraging the DINOv2 framework, we train Vision Transformers on automatically mined face crops, eliminating the need for identity labels. Our method demonstrates strong open-set re-identification performance, surpassing supervised baselines on challenging benchmarks such as Bossou, despite utilising no labelled data during training. This work underscores the potential of self-supervised learning in biodiversity monitoring and paves the way for scalable, non-invasive population studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。