无需训练即可追踪野外多种动物,跨物种跨环境表现优异。
Zero-Shot Multi-Animal Tracking in the Wild
- 用SAM2和Grounding DINO构建零样本追踪框架,不需调参或重训。
- 在4个数据集上达到最新水平,覆盖灵长类、鸟类等多类动物。
- 适合野外复杂场景下无标注数据的动物行为研究者使用。
多动物追踪对理解动物生态与行为至关重要,但受栖息地差异、运动模式多样及物种外观变化影响,仍具挑战。传统方法通常需针对每个新场景大量微调与启发式设计。本文探索视觉基础模型在零样本多动物追踪中的应用。基于SAM2MOT,我们将Grounding DINO与分割一切模型2(SAM 2)结合,并引入三项针对性改进,使框架能适应动物外观与行为,且无需在不同数据集间重新训练或调整超参数。我们还评估了近期的SAM3模型,发现其在野外多动物追踪中存在实际局限性。所提方法在Chimp-Act、Bird Flock Tracking、AnimalTrack及GMOT-40子集上均达当前最佳性能,展现出强大的跨物种与跨环境泛化能力。代码已公开于https://github.com/ecker-lab/SAM2-Animal-Tracking。
原文摘要 · Abstract (English)
Multi-animal tracking is crucial for understanding animal ecology and behavior, yet remains challenging due to variations in habitat, motion patterns, and species appearance. Traditional approaches typically require extensive fine-tuning and heuristic design for each new scenario. In this work, we explore vision foundation models for zero-shot multi-animal tracking. Building on SAM2MOT, we combine Grounding DINO with the Segment Anything Model2 (SAM 2) and introduce three targeted modifications to adapt the framework to animal appearance and behavior without any retraining or hyperparameter tuning between datasets. We also evaluate the recent SAM3 model, but identify practical limitations that restrict its applicability to multi-animal tracking in the wild. Our method achieves state-of-the-art results across Chimp-Act, Bird Flock Tracking, AnimalTrack, and a subset of GMOT-40, demonstrating robust generalization across diverse species and environments. The code is available at https://github.com/ecker-lab/SAM2-Animal-Tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。