arXiv:2501.05098cs.CV2025-01被引 45

构建大规模多模态人体运动数据集,支持表情、手势等精细动作描述。

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset

  • 基于自动标注流程从视频中提取3D全身动作与纹理标签。
  • 包含1950万帧全身姿态标注,覆盖12万段动作序列和80.8万视频。
  • 适合做文本/音频驱动的动作生成、3D人体网格重建等研究。

本文提出Motion-X++,一个大规模多模态3D全身人体动作数据集。现有数据集多仅捕捉身体姿态,缺乏面部表情、手部动作及细粒度姿态描述,且通常局限于实验室环境,依赖人工文本标注,难以扩展。为此,我们设计可扩展的标注流程,能从RGB视频中自动获取3D全身动作与全面纹理标签,构建出包含81.1万组文本-动作对的Motion-X数据集。进一步通过优化标注流程、引入更多模态并扩大数据量,形成Motion-X++,包含1950万帧3D全身姿态标注、12.05万段动作序列、80.8万段RGB视频、45.3万段音频、1950万帧级全身姿态描述及12.05万序列级语义标签。大量实验验证了标注流程的准确性,并展示了Motion-X++在文本驱动、音频驱动的全身动作生成、3D全身人体网格恢复、2D全身关键点估计等下游任务中的显著优势。

原文摘要 · Abstract (English)

In this paper, we introduce Motion-X++, a large-scale multimodal 3D expressive whole-body human motion dataset. Existing motion datasets predominantly capture body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions, and are typically limited to lab settings with manually labeled text descriptions, thereby restricting their scalability. To address this issue, we develop a scalable annotation pipeline that can automatically capture 3D whole-body human motion and comprehensive textural labels from RGB videos and build the Motion-X dataset comprising 81.1K text-motion pairs. Furthermore, we extend Motion-X into Motion-X++ by improving the annotation pipeline, introducing more data modalities, and scaling up the data quantities. Motion-X++ provides 19.5M 3D whole-body pose annotations covering 120.5K motion sequences from massive scenes, 80.8K RGB videos, 45.3K audios, 19.5M frame-level whole-body pose descriptions, and 120.5K sequence-level semantic labels. Comprehensive experiments validate the accuracy of our annotation pipeline and highlight Motion-X++'s significant benefits for generating expressive, precise, and natural motion with paired multimodal labels supporting several downstream tasks, including text-driven whole-body motion generation,audio-driven motion generation, 3D whole-body human mesh recovery, and 2D whole-body keypoints estimation, etc.

动作数据集多模态3D人体建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。