用自监督方法提升灵长类行为识别,无需标注数据就能显著提效。
Domain-Adaptive Pretraining Improves Primate Behavior Recognition
- 在域适应预训练中继续使用灵长类视频数据预训练模型。
- 在两个数据集上准确率和mAP分别提升6.1%和6.3个百分点。
- 仅需无标签数据即可提升性能,适合缺乏标注的动物行为研究。
动物行为的计算机视觉分析为生态学、认知研究和保护工作提供了有力工具。视频陷阱可大规模采集数据,但标注成本仍是构建大规模数据集的瓶颈。因此需要数据高效的学习方法。本文展示通过自监督学习可显著提升灵长类行为的动作识别性能。在两个大猩猩行为数据集(PanAf 和 ChimpACT)上,我们分别比已有最优模型提高6.1%准确率和6.3% mAP。该成果得益于使用预训练的 V-JEPA 模型并实施领域自适应预训练(DAP),即在本领域数据上继续预训练。我们证实性能提升主要来自 DAP。该方法无需标注样本,具有极大潜力提升动物行为识别能力。代码已公开于 https://github.com/ecker-lab/dap-behavior。
原文摘要 · Abstract (English)
Computer vision for animal behavior offers promising tools to aid research in ecology, cognition, and to support conservation efforts. Video camera traps allow for large-scale data collection, but high labeling costs remain a bottleneck to creating large-scale datasets. We thus need data-efficient learning approaches. In this work, we show that we can utilize self-supervised learning to considerably improve action recognition on primate behavior. On two datasets of great ape behavior (PanAf and ChimpACT), we outperform published state-of-the-art action recognition models by 6.1 %pt. accuracy and 6.3 %pt. mAP, respectively. We achieve this by utilizing a pretrained V-JEPA model and applying domain-adaptive pretraining (DAP), i.e. continuing the pretraining with in-domain data. We show that most of the performance gain stems from the DAP. Our method promises great potential for improving the recognition of animal behavior, as DAP does not require labeled samples. Code is available at https://github.com/ecker-lab/dap-behavior
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。