从YouTube自动收集30万段动物视频,构建首个4D动物重建基准
Web-Scale Collection of Video Data for 4D Animal Reconstruction
- 自动化爬取并处理YouTube视频,生成以动物为中心的剪辑
- 建立包含11000帧的AiM基准数据集,覆盖多样动物动作
- 提出首个无模型4D动物重建基线,解决评估体系不匹配问题
动物计算机视觉对野生动物研究前景广阔,但依赖大规模数据,而现有采集方法多基于受控环境。近期数据驱动方法展示了单视角、非侵入式分析潜力,但当前动物视频数据集规模有限——仅提供约2.4K个15帧片段,且缺乏针对动物中心3D/4D任务的关键处理。本文提出自动化管道,从YouTube挖掘视频并处理为动物中心剪辑,附带姿态估计、追踪和3D/4D重建所需的辅助标注。利用该管道,我们收集了30,000段视频(200万帧),较之前工作提升一个数量级。为验证其价值,聚焦四维四足动物重建任务,构建了人工筛选的动物动态(AiM)基准,包含230个序列、11,000帧,展示清晰多样的动物运动。我们在AiM上评估了先进模型驱动与无模型方法,发现2D指标偏好前者,尽管其3D形状失真;后者重建更自然但得分较低,暴露出当前评估体系的差距。为此,我们对最新无模型方法引入序列级优化,建立首个4D动物重建基线。本工作提供的管道、基准与基线旨在推动大规模、无标记4D动物重建及相关任务从野外视频中发展。代码与数据集见:https://github.com/briannlongzhao/Animal-in-Motion。
原文摘要 · Abstract (English)
Computer vision for animals holds great promise for wildlife research but often depends on large-scale data, while existing collection methods rely on controlled capture setups. Recent data-driven approaches show the potential of single-view, non-invasive analysis, yet current animal video datasets are limited--offering as few as 2.4K 15-frame clips and lacking key processing for animal-centric 3D/4D tasks. We introduce an automated pipeline that mines YouTube videos and processes them into object-centric clips, along with auxiliary annotations valuable for downstream tasks like pose estimation, tracking, and 3D/4D reconstruction. Using this pipeline, we amass 30K videos (2M frames)--an order of magnitude more than prior works. To demonstrate its utility, we focus on the 4D quadruped animal reconstruction task. To support this task, we present Animal-in-Motion (AiM), a benchmark of 230 manually filtered sequences with 11K frames showcasing clean, diverse animal motions. We evaluate state-of-the-art model-based and model-free methods on Animal-in-Motion, finding that 2D metrics favor the former despite unrealistic 3D shapes, while the latter yields more natural reconstructions but scores lower--revealing a gap in current evaluation. To address this, we enhance a recent model-free approach with sequence-level optimization, establishing the first 4D animal reconstruction baseline. Together, our pipeline, benchmark, and baseline aim to advance large-scale, markerless 4D animal reconstruction and related tasks from in-the-wild videos. Code and datasets are available at https://github.com/briannlongzhao/Animal-in-Motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。