arXiv:2607.04518cs.CV2026-07

用视频分析动物步态,实现不接触的个体识别。

A non-invasive video-based method for individual identification of wildlife using gait dynamics

论文配图:A non-invasive video-based method for individual identification of wildlife using gait dynamics
图 1 · 摘自论文原文
  • 通过分割模型和双分支网络提取步态时空特征。
  • 跨物种数据验证显示个体间步态差异显著,相似度分布清晰。
  • 适合野外生态监测,无需标记或捕获动物。

步态是独特的行为特征,可实现无需物理接触的野生动物个体识别。尽管人类步态分析已广泛研究,但受环境变化和缺乏可扩展方法限制,其在野生动物中的应用仍有限。本文提出一种全自动视频驱动的步态分析与个体识别流水线,采用深度时空表征学习。该方法利用 Segment Anything Model 3 (SAM3) 生成高质量的 RGB 和二值轮廓掩码,有效从复杂自然背景中分离动物。分割后的视频序列由 ResNet18 提取空间特征,VideoPrism(基于变压器的视频模型)建模时间运动特征。两模型经分类任务微调后作为特征提取器,生成具有区分性的步态表征。通过余弦相似度比较步态签名,实现无需物理标记的个体相似性聚类。多源、跨物种视频实验表明,个体内部步态高度一致,个体间差异明显。定量结果基于余弦相似度分布与轮廓得分,验证了方法有效性。结果表明,步态动态为野生动物个体识别提供了一种可行的非侵入性方案,并凸显视频驱动深度学习流水线在可扩展生态监测中的潜力。

原文摘要 · Abstract (English)

Gait is a distinctive behavioral characteristic that enables non-invasive individual identification without requiring physical interaction with an animal. While gait-based analysis has been extensively studied in humans, its application to wildlife remains limited due to environmental variability and the lack of scalable identification methods. This paper presents a fully automated, video-based pipeline for wildlife gait analysis and individual identification using deep spatiotemporal representation learning. The proposed pipeline uses the Segment Anything Model 3 (SAM3) to generate high-quality RGB and binary silhouette masks, robustly isolating animals from complex natural backgrounds. Segmented video sequences are processed using a convolutional neural network (ResNet18) for spatial feature extraction and a transformer-based video model (VideoPrism) for temporal motion modeling. Both models are fine-tuned using a classification objective and subsequently used as feature extractors to generate discriminative gait representations. Cosine similarity is then used to compare gait signatures, enabling similarity-based clustering of individuals without reliance on physical markings or invasive tagging. Experiments conducted on multi-source wildlife video data across multiple species demonstrate strong intra-individual consistency and clear inter-individual separation. Quantitative results using cosine similarity distributions and silhouette scores confirm the effectiveness of the proposed method. These findings demonstrate that gait dynamics provide a viable, non-invasive approach for individual identification in wildlife and highlight the potential of video-based deep learning pipelines for scalable ecological monitoring.

步态识别野生动物视频分析非侵入式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。