用网络视频自动构建大规模真实人体动作数据集
XmoPipe: A Pipeline for Large-Scale In-the-Wild Human Motion Dataset Construction

- 关键词检索视频,提取3D身体与面部动作
- 生成的动作用于训练,性能媲美传统采集数据
- 适合需要多样化真实动作数据的研究者
大规模人体动作数据集对分析、合成和理解人体运动模型至关重要。虽然标记式动作捕捉能提供精确数据,但成本高、规模和多样性有限。近年来单目动作捕捉与视频-语言理解的进步,使得从非受限网络视频中提取合理动作成为可能。我们提出一个可扩展的在野人体动作数据集构建流水线:通过少量关键词检索视频,提取3D身体与面部运动,并生成高层文本描述。该流程灵活,可针对性收集多种动作、多人交互或表达性行为。通过训练动作重建与生成模型验证其质量,结果表明性能可比肩基于传统动作捕捉数据训练的模型,且具备强跨数据集泛化能力。
原文摘要 · Abstract (English)
Large-scale human motion datasets are essential for training robust motion models for analysis, synthesis, and understanding. While marker-based motion capture provides precise data, it is costly and limited in scale and diversity. Recent advances in monocular motion capture and video-language understanding open the way to extract plausible motion from unconstrained online videos. We present a scalable pipeline for constructing in-the-wild human motion datasets. From a few keywords, the system retrieves videos, extracts 3D body and facial motion, and generates high-level textual descriptions. The pipeline is flexible, enabling targeted collection of various motions, multi-person interactions, or expressive behaviors. We demonstrate its quality by training motion reconstruction and motion generation models, showing performance comparable to models trained on traditional motion capture datasets and strong cross-dataset generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。