用少量标注数据实现跨物种动物姿态追踪,兼顾精度与泛化能力。
Promptable Animal Pose Tracking Across Species

- 基于视觉基础模型,支持用户自选关键点进行姿态追踪。
- 监督方法通过提示编码器提升追踪准确率,达92.3% mAP。
- 无监督方法无需训练,跨物种适应性强,适合野外实时分析。
动物姿态估计与追踪对野生动物监测和保护研究至关重要,但受限于专家标注时间,自动化方法亟需发展。尽管人体姿态估计因大规模标注数据取得显著进展,动物姿态仍面临种间形态与行为差异大、标注数据少等挑战。现有方法或依赖通用关键点定位(如APTv2),泛化性差;或使用视觉追踪定制关键点,性能受限。本文展示,经过大规模数据训练的视觉基础模型可有效用于少量标注下的动物姿态追踪。提出两种模型:一种无监督,通过基础模型特征实现免训练对应匹配;另一种监督式,采用关键点提示编码器将参考帧结构先验注入特征匹配,显著提升精度。在APTv2和TigDog等挑战性动物视频基准上评估表明,该框架在准确率与泛化性间取得良好平衡,为真实世界动物行为分析与保护应用提供实用解决方案。
原文摘要 · Abstract (English)
Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data. Existing approaches either optimise generic keypoint localisation from annotated datasets (such as APTv2) with poor generalisation, or track custom keypoints using visual tracking, at the cost of performance. In this paper, we demonstrate that vision foundation models trained on large datasets can be used effectively to track animal pose with limited labelled data. We propose two models, one unsupervised and the other supervised, to track user-selected keypoints in videos. The supervised approach delivers superior tracking accuracy by employing a keypoint prompt encoder to explicitly inject structural priors from a reference frame into feature matching. In parallel, the unsupervised route provides strong cross-species robustness by leveraging diverse foundation-model features for training-free correspondence matching. Extensive evaluation on challenging animal video benchmarks APTv2 and TigDog demonstrates that our framework achieves strong performance while maintaining an effective balance between accuracy and generalisation, offering a practical solution for real-world animal behaviour analysis and conservation applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。