构建首个大规模多视角RGB-D犬类4D运动数据集,推动动物动作重建研究。
DogMo: A Large-Scale Multi-View RGB-D Dataset for 4D Canine Motion Recovery
- 采集10只犬、1200段动作的多视角RGB-D视频,覆盖多样品种与行为。
- 提出三阶段优化流程,基于SMAL模型实现高精度犬体形与姿态重建。
- 支持单目与多视角、RGB与RGB-D输入的系统性评估,适合跨领域研究者。
我们提出DogMo,一个大规模多视角RGB-D视频数据集,用于从图像中恢复犬类4D运动。该数据集包含10只不同犬种的1200段运动序列,涵盖丰富多样的动作与品种差异,解决了现有数据集缺乏多视角真实3D信息、规模与多样性不足的问题。基于DogMo,我们建立了四个运动恢复基准任务,支持单目与多视角、RGB与RGB-D输入的系统评估。为实现精准运动重建,我们进一步提出一种三阶段实例特定优化流程,通过粗对齐、密集对应监督与时间正则化逐步优化SMAL模型的体形与姿态。本数据集与方法为犬类运动恢复研究提供了坚实基础,并拓展了计算机视觉、图形学与动物行为建模交叉领域的研究方向。
原文摘要 · Abstract (English)
We present DogMo, a large-scale multi-view RGB-D video dataset capturing diverse canine movements for the task of motion recovery from images. DogMo comprises 1.2k motion sequences collected from 10 unique dogs, offering rich variation in both motion and breed. It addresses key limitations of existing dog motion datasets, including the lack of multi-view and real 3D data, as well as limited scale and diversity. Leveraging DogMo, we establish four motion recovery benchmark settings that support systematic evaluation across monocular and multi-view, RGB and RGB-D inputs. To facilitate accurate motion recovery, we further introduce a three-stage, instance-specific optimization pipeline that fits the SMAL model to the motion sequences. Our method progressively refines body shape and pose through coarse alignment, dense correspondence supervision, and temporal regularization. Our dataset and method provide a principled foundation for advancing research in dog motion recovery and open up new directions at the intersection of computer vision, computer graphics, and animal behavior modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。