arXiv:2412.00837cs.CV2024-12CVPR被引 25

用家族感知Transformer提升多种四足动物姿态与形状估计精度

AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer

论文配图:AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer
图 1 · 摘自论文原文
  • 采用高容量Transformer与家族监督对比学习统一建模不同四足动物
  • 在41.3k张标注图像上训练,3D估计精度超越现有方法
  • 适合动物行为学、生物力学研究者使用,尤其关注跨物种分析

动物行为与生物力学的定量分析依赖于跨物种的精准姿态与形状估计,对动物福利与生物学研究至关重要。然而,以往方法网络容量有限且多物种数据集稀缺,限制了该领域发展。本文提出AniMer,基于家族感知Transformer实现多种四足动物的姿态与形状估计,显著提升重建精度。核心思路是融合高容量Transformer骨干网络与动物家族监督对比学习,统一建模不同四足类别的判别性特征。为支持有效训练,整合现有开源四足动物数据集(含2D/3D标注),共获41.3k张标注图像。为扩充3D数据多样性,提出新生成式数据集CtrlAni3D,基于扩散模型条件图像生成流程,包含约10,000张带像素对齐SMAL标签的图像。实验表明,AniMer在Animal3D、CtrlAni3D等3D数据集及分布外的Animal Kingdom数据集上均优于现有方法。消融实验进一步验证了网络设计与CtrlAni3D的有效性,适用于真实场景应用。

原文摘要 · Abstract (English)

Quantitative analysis of animal behavior and biomechanics requires accurate animal pose and shape estimation across species, and is important for animal welfare and biological research. However, the small network capacity of previous methods and limited multi-species dataset leave this problem underexplored. To this end, this paper presents AniMer to estimate animal pose and shape using family aware Transformer, enhancing the reconstruction accuracy of diverse quadrupedal families. A key insight of AniMer is its integration of a high-capacity Transformer-based backbone and an animal family supervised contrastive learning scheme, unifying the discriminative understanding of various quadrupedal shapes within a single framework. For effective training, we aggregate most available open-sourced quadrupedal datasets, either with 3D or 2D labels. To improve the diversity of 3D labeled data, we introduce CtrlAni3D, a novel large-scale synthetic dataset created through a new diffusion-based conditional image generation pipeline. CtrlAni3D consists of about 10k images with pixel-aligned SMAL labels. In total, we obtain 41.3k annotated images for training and validation. Consequently, the combination of a family aware Transformer network and an expansive dataset enables AniMer to outperform existing methods not only on 3D datasets like Animal3D and CtrlAni3D, but also on out-of-distribution Animal Kingdom dataset. Ablation studies further demonstrate the effectiveness of our network design and CtrlAni3D in enhancing the performance of AniMer for in-the-wild applications. The project page of AniMer is https://luoxue-star.github.io/AniMer_project_page/.

姿态估计四足动物Transformer合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。