构建大规模高质量人体动作数据集,提升文本生成动作的精度与多样性。
RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation

- 设计分层语义分类体系,自动过滤静态和低质动作序列。
- 模型在RoMo上训练后,动作真实度与多样性均达当前最优水平。
- 适合关注动作生成、人机交互或具身智能的研究者使用。
语言、图像和视频生成的成功表明,大规模精心标注的数据集是构建强大生成模型的关键。然而,3D人体动作生成仍落后,受限于小规模高保真动捕数据集与大规模但包含大量静态或低质量序列的野外数据集之间的权衡。我们提出RoMo,一个大规模、高质量、精心组织的野外人体动作数据集,解决了这一矛盾。为保证质量,引入基于语义分类体系的过滤流程,主动剔除静态及易出错序列。每段动作均配有详细描述,并按创新的三级语义分类体系组织。该层级结构支持细粒度的类别级评估,揭示全局指标掩盖的模型优缺点。实验表明,在RoMo上训练的模型在动作保真度与多样性方面达到当前最佳表现,同时能更准确理解复杂细微的文本指令。最后,我们发布运动工具箱(Motion Toolbox),统一评估指标、数据转换与可视化流程,为可复现、可解释的动作生成研究奠定基础。
原文摘要 · Abstract (English)
Success in generative modeling across language, image, and video demonstrates that large, well-curated datasets are the key driver for building capable models. 3D Human motion, however, has lagged behind, constrained by an unsatisfying choice between small, high-fidelity motion capture datasets and large-scale in-the-wild collections dominated by static or low-quality sequences. We introduce RoMo, a rich, large-scale, carefully curated dataset of in-the-wild human motions that resolves these tradeoffs. To ensure quality, we introduce a taxonomy-aware filtering pipeline that aggressively removes static and artifact-prone sequences. Every sequence is annotated with detailed captions and organized by a novel three-level semantic taxonomy. This hierarchical structure enables fine-grained, per-category evaluation, that reveals model strengths and weaknesses obscured by global metrics. We demonstrate that models trained on RoMo achieve state-of-the-art fidelity and diversity while gaining a superior understanding of complex, subtle text prompts. Finally, we release the Motion Toolbox to standardize metrics, data conversion, and visualization, establishing a foundation for reproducible and interpretable motion generation research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。