用扩散模型自动标注动物关键点,提升跨物种识别准确率
Categorical Keypoint Positional Embedding for Robust Animal Re-Identification
- 用单张标注图+预训练扩散模型自动传播关键点,大幅降低人工标注成本
- 引入类别化关键点位置嵌入,使ViT学习更鲁棒的语义特征
- 在4个野生动物数据集上超越现有方法,适合生态监测场景
动物重识别(ReID)已成为生态研究中不可或缺的工具,在追踪种群动态、分析行为模式和评估生态影响方面发挥关键作用,对制定科学保护策略至关重要。与人体ReID不同,动物ReID因姿态变化大、环境差异显著,且无法直接应用预训练模型于动物数据,导致跨物种识别更加复杂。本文提出一种创新的关键点传播机制,仅需一张标注图像和预训练扩散模型,即可将关键点传播至整个数据集,显著降低人工标注成本。同时,通过引入关键点位置编码(KPE)和类别化关键点位置嵌入(CKPE),增强视觉变换器(ViT)的表达能力,使其学习更鲁棒且具语义感知的特征表示。该方法提供更全面、细致的关键点表征,实现更高精度和效率的重识别。大量实验证明,该方法在四个野生动物数据集上显著优于现有最先进方法。代码将公开发布。
原文摘要 · Abstract (English)
Animal re-identification (ReID) has become an indispensable tool in ecological research, playing a critical role in tracking population dynamics, analyzing behavioral patterns, and assessing ecological impacts, all of which are vital for informed conservation strategies. Unlike human ReID, animal ReID faces significant challenges due to the high variability in animal poses, diverse environmental conditions, and the inability to directly apply pre-trained models to animal data, making the identification process across species more complex. This work introduces an innovative keypoint propagation mechanism, which utilizes a single annotated image and a pre-trained diffusion model to propagate keypoints across an entire dataset, significantly reducing the cost of manual annotation. Additionally, we enhance the Vision Transformer (ViT) by implementing Keypoint Positional Encoding (KPE) and Categorical Keypoint Positional Embedding (CKPE), enabling the ViT to learn more robust and semantically-aware representations. This provides more comprehensive and detailed keypoint representations, leading to more accurate and efficient re-identification. Our extensive experimental evaluations demonstrate that this approach significantly outperforms existing state-of-the-art methods across four wildlife datasets. The code will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。