用灵长类动物视频数据训练通用模型,提升小样本下的行为识别能力。
PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- 以灵长类动物为中心构建大规模视频预训练数据集PriVi,实现数据驱动的模型训练。
- 在4个基准数据集上优于现有方法,仅用少量标注数据即达高精度。
- 适合从事灵长类行为分析、低资源场景下视频理解的研究者使用。
非人灵长类是人类最近的现存亲属,研究其行为对认知、进化与保护领域至关重要。计算机视觉可极大助力该研究,但现有方法多依赖以人为中心的预训练模型且局限于单一数据集,限制了泛化能力。本文提出从模型主导转向数据主导,构建大规模灵长类为中心的视频预训练数据集PriVi,包含424小时精选视频:174小时来自11个研究场景的行为数据,250小时通过可扩展的数据清洗流程获取的网络视频。在此基础上,继续预训练V-JEPA大型视频模型,学习灵长类特异性表征,并采用轻量级冻结分类器评估。在ChimpACT、PanAf500、BaboonLand和ChimpBehave四个基准数据集上,该方法持续优于先前工作,包括全微调基线,且标签数量较少时仍表现良好。首次证明视频模型中‘领域级预训练’(在相似数据上而非目标数据本身上预训练)有效。灵长类中心预训练显著提升数据效率与泛化性能,为低标签应用提供可行方案。数据集、代码与模型已公开:https://privi.eckerlab.org
原文摘要 · Abstract (English)
Non-human primates are our closest living relatives, and analyzing their behavior is central to research in cognition, evolution, and conservation. Computer vision could greatly aid this research, but existing methods often rely on human-centric pretrained models and focus on single datasets, which limits generalization. We address this limitation by shifting from a model-centric to a data-centric approach and introduce PriVi, a large-scale primate-centric video pretraining dataset. PriVi contains 424 hours of curated video, combining 174 hours from behavioral research across 11 settings with 250 hours of diverse web-sourced footage, assembled through a scalable data curation pipeline. We continue pretraining V-JEPA, a large-scale video model, on PriVi to learn primate-specific representations and evaluate it using a lightweight frozen classifier. Across four benchmark datasets, ChimpACT, PanAf500, BaboonLand, and ChimpBehave, our approach consistently outperforms prior work, including fully finetuned baselines, and scales favorably with fewer labels. These results demonstrate for the first time that domain-level pretraining, where pretraining is conducted on similar data but not the target dataset itself, works for video models. Our primate-centric pretraining substantially improves data efficiency and generalization, making it a promising approach for low-label applications. Dataset, code, and models are available: https://privi.eckerlab.org
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。