arXiv:2412.08101cs.CVcs.LG2024-12ICCV被引 12

用生成模型合成百万张动物3D姿态图像,训练出顶尖的动物姿态估计器。

Generative Zoo

  • 用条件图像生成模型合成逼真动物图像及对应3D参数
  • 在真实数据集上达到当前最佳性能,仅用合成数据训练
  • 适合动物行为建模、计算机视觉研究者使用

基于模型的3D动物姿态与形状估计使动物行为的计算建模成为可能。但训练此类模型需要大量带有精确姿态与形状标注的图像数据,而获取这些数据需依赖多视角或标记式动作捕捉系统,难以适用于野外动物且无法扩展到众多物种。部分研究通过人工2D标注后优化3D参数来伪标注真实图像,但因单目重建问题的病态性,得到的姿态与形状常不自然。另一些工作采用游戏引擎和艺术家设计的3D资产生成合成数据,虽有完美标注但缺乏视觉真实感且适配新物种耗时。为此,我们提出一种替代方案:利用条件图像生成模型进行渲染。我们构建了一个流程,可为多种哺乳类四足动物采样多样姿态与形状,并生成带有真实3D参数的逼真图像。为验证可扩展性,我们推出GenZoo,一个包含一百万张不同主体图像的合成数据集。我们在GenZoo上训练3D姿态与形状回归器,在真实世界动物姿态与形状估计基准测试中取得当前最优表现,且训练全程仅使用合成数据。

原文摘要 · Abstract (English)

The model-based estimation of 3D animal pose and shape from images enables computational modeling of animal behavior. Training models for this purpose requires large amounts of labeled image data with precise pose and shape annotations. However, capturing such data requires the use of multi-view or marker-based motion-capture systems, which are impractical to adapt to wild animals in situ and impossible to scale across a comprehensive set of animal species. Some have attempted to address the challenge of procuring training data by pseudo-labeling individual real-world images through manual 2D annotation, followed by 3D-parameter optimization to those labels. While this approach may produce silhouette-aligned samples, the obtained pose and shape parameters are often implausible due to the ill-posed nature of the monocular fitting problem. Sidestepping real-world ambiguity, others have designed complex synthetic-data-generation pipelines leveraging video-game engines and collections of artist-designed 3D assets. Such engines yield perfect ground-truth annotations but are often lacking in visual realism and require considerable manual effort to adapt to new species or environments. Motivated by these shortcomings, we propose an alternative approach to synthetic-data generation: rendering with a conditional image-generation model. We introduce a pipeline that samples a diverse set of poses and shapes for a variety of mammalian quadrupeds and generates realistic images with corresponding ground-truth pose and shape parameters. To demonstrate the scalability of our approach, we introduce GenZoo, a synthetic dataset containing one million images of distinct subjects. We train a 3D pose and shape regressor on GenZoo, which achieves state-of-the-art performance on a real-world animal pose and shape estimation benchmark, despite being trained solely on synthetic data. https://genzoo.is.tue.mpg.de

3D姿态估计生成模型合成数据动物行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。