统一估计哺乳动物与鸟类姿态形状,提升跨物种空间智能
AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware Transformer
- 用家族感知的视觉变换器+专家混合架构,分层学习共性与特异性特征
- 合成12.4万张鸟的3D图像数据集,解决鸟类真实数据稀缺问题
- 在跨物种挑战数据集上表现领先,适合生物研究与多物种智能系统
在基础模型时代,通过单一网络实现对不同动态物体的统一理解,有望增强空间智能。准确估计跨多种类动物的姿态与形状,对生物研究中的定量分析至关重要。然而,受限于以往方法的网络容量及全面多物种数据集的缺乏,该领域仍被忽视。为此,我们提出AniMer+,即可扩展AniMer框架的升级版。本文聚焦于哺乳动物(mammalia)与鸟类(aves)的统一姿态与形状重建。AniMer+的核心创新在于高容量、家族感知的视觉变换器(ViT),结合专家混合(MoE)设计,将网络层划分为物种特异组件(哺乳类与鸟类)和共享组件,从而高效学习共性与差异解剖特征。为弥补3D训练数据严重不足,尤其是鸟类数据,我们引入基于扩散的条件图像生成流程,构建两个大规模合成数据集:用于四足动物的CtrlAni3D与用于鸟类的CtrlAVES3D。其中,CtrlAVES3D是首个大规模、带3D标注的鸟类数据集,对解决单视角深度歧义至关重要。模型在41.3k张哺乳动物与12.4k张鸟类图像(含真实与合成数据)上联合训练,广泛基准测试中超越现有方法,包括具有挑战性的域外动物王国数据集。消融实验验证了新架构与合成数据集对实际应用性能的提升作用。
原文摘要 · Abstract (English)
In the era of foundation models, achieving a unified understanding of different dynamic objects through a single network has the potential to empower stronger spatial intelligence. Moreover, accurate estimation of animal pose and shape across diverse species is essential for quantitative analysis in biological research. However, this topic remains underexplored due to the limited network capacity of previous methods and the scarcity of comprehensive multi-species datasets. To address these limitations, we introduce AniMer+, an extended version of our scalable AniMer framework. In this paper, we focus on a unified approach for reconstructing mammals (mammalia) and birds (aves). A key innovation of AniMer+ is its high-capacity, family-aware Vision Transformer (ViT) incorporating a Mixture-of-Experts (MoE) design. Its architecture partitions network layers into taxa-specific components (for mammalia and aves) and taxa-shared components, enabling efficient learning of both distinct and common anatomical features within a single model. To overcome the critical shortage of 3D training data, especially for birds, we introduce a diffusion-based conditional image generation pipeline. This pipeline produces two large-scale synthetic datasets: CtrlAni3D for quadrupeds and CtrlAVES3D for birds. To note, CtrlAVES3D is the first large-scale, 3D-annotated dataset for birds, which is crucial for resolving single-view depth ambiguities. Trained on an aggregated collection of 41.3k mammalian and 12.4k avian images (combining real and synthetic data), our method demonstrates superior performance over existing approaches across a wide range of benchmarks, including the challenging out-of-domain Animal Kingdom dataset. Ablation studies confirm the effectiveness of both our novel network architecture and the generated synthetic datasets in enhancing real-world application performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。